Tensor Core systems and hardware chips for large language models

By designing a tensor core system for large language models, including activation matrix processing components, weight matrix processing components, tensor core computing components, accumulation matrix and result processing components, the problems of low computing power utilization and waste of bandwidth in large language models are solved, and more efficient computing and lower power consumption are achieved.

CN119576273BActive Publication Date: 2025-05-13AEROSPACE INFORMATION RES INST CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510143463.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-13
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

When used in large language models, existing hardware inference chips have problems with bandwidth and power consumption waste caused by low computing power utilization and complex data transmission, and the design architecture cannot meet the digital type and computing requirements of large language models.

Method used

Design a tensor core system for large language models, including activation matrix processing components, weight matrix processing components, tensor core operation components, accumulation matrix and result processing components, through which data transposition preprocessing and subsequent multiplication operations are independently completed, saving bandwidth, reducing chip area and power consumption.

Benefits of technology

It improves the chip computing power utilization rate, reduces the waste of bandwidth and power consumption, makes data flow more smooth and efficient, and meets the computing needs of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576273B_ABST
    Figure CN119576273B_ABST
Patent Text Reader

Abstract

The present invention provides a tensor core system and hardware chip for a large language model, which relates to the field of artificial intelligence chip technology, including: an activation matrix processing component, which is used to cache and preprocess the activation matrix data read from the memory to obtain a target activation matrix; a weight matrix processing component, which is used to cache and preprocess the weight matrix data read from the memory to obtain a target weight matrix; a tensor core operation component, which is used to perform matrix floating-point multiplication and addition operations on the target activation matrix and the target weight matrix to obtain a floating-point multiplication and addition operation result, and based on a preset floating-point number system, the accumulation matrix and the floating-point multiplication and addition operation result are accumulated to obtain a target result matrix; an accumulation matrix and result processing component, which is used to perform floating-point number system conversion processing on the target result matrix to obtain a result matrix to be written, and write the result matrix to be written into the memory. The present invention is more conducive to improving the utilization rate of chip computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence chip technology, and in particular to a tensor core system and hardware chip for a large language model. Background Art

[0002] In the case of large language models, there are currently no hardware inference chips customized for large language models.

[0003] The hardware inference chips currently used in large language models are mainly developed and manufactured for general scenarios. Their design architectures cannot meet the number types and computing requirements required by large language models. As a result, when these hardware inference chips are applied to large language models, there are problems such as low chip computing power utilization and complex data transmission process leading to waste of bandwidth and power consumption.

[0004] Therefore, there is an urgent need for a tensor core system and hardware chip for large language models to solve the above problems. Summary of the invention

[0005] In view of the problems existing in the prior art, the present invention provides a tensor core system and a hardware chip for a large language model.

[0006] The present invention provides a tensor core system for a large language model, comprising:

[0007] An activation matrix processing component is used to cache the activation matrix data read from the memory, and perform activation matrix preprocessing on the cached activation matrix data to obtain a target activation matrix;

[0008] A weight matrix processing component is used to cache the weight matrix data read from the memory, and perform weight matrix preprocessing on the cached weight matrix data to obtain a target weight matrix;

[0009] A tensor core operation component is used to perform matrix floating-point multiplication and addition operations on the target activation matrix and the target weight matrix to obtain a floating-point multiplication and addition operation result, and to accumulate the accumulation matrix stored in the memory and the floating-point multiplication and addition operation result based on a preset floating-point number system to obtain a target result matrix;

[0010] The accumulation matrix and result processing component is used to perform floating point number conversion processing on the target result matrix to obtain the result matrix to be written, and write the result matrix to be written into the memory.

[0011] According to a tensor core system for a large language model provided by the present invention, the activation matrix processing component includes an activation matrix AXI bus control unit, an activation matrix buffer unit and an activation matrix preprocessing unit, wherein:

[0012] The activation matrix AXI bus control unit is used to determine the activation matrix storage address information corresponding to the activation matrix data in the memory according to the dimension information of the activation matrix data, and send an activation matrix read request to the memory, wherein the activation matrix read request is constructed according to the activation matrix storage address information;

[0013] The activation matrix buffer unit is used to cache the activation matrix data;

[0014] The activation matrix preprocessing unit is used to perform matrix transposition on the activation matrix data currently cached in the activation matrix cache unit, and based on the floating-point number system of the activation matrix data currently cached in the activation matrix cache unit, perform corresponding decompression processing on the floating-point elements in the activation matrix data after the matrix transposition to obtain the target activation matrix.

[0015] According to a tensor core system for a large language model provided by the present invention, the weight matrix processing component includes a weight matrix AXI bus control unit, a weight matrix buffer unit and a weight matrix preprocessing unit, wherein:

[0016] The weight matrix AXI bus control unit is used to determine the weight matrix storage address information corresponding to the weight matrix data in the memory according to the dimension information and quantization information of the weight matrix data, and send a weight matrix read request to the memory, wherein the weight matrix read request is constructed according to the weight matrix storage address information;

[0017] The weight matrix buffer unit is used to cache the weight matrix data;

[0018] The weight matrix preprocessing unit is used to perform matrix transposition on the weight matrix data currently cached in the weight matrix cache area unit, and based on the floating-point number system of the weight matrix data currently cached in the weight matrix cache area unit, perform corresponding decompression processing on the floating-point elements in the weight matrix data after the matrix transposition to obtain the target weight matrix.

[0019] According to a tensor core system for a large language model provided by the present invention, the activation matrix preprocessing unit is further used to perform a corresponding matrix transposition on the activation matrix data currently cached to the activation matrix cache unit based on a preset matrix multiplication instruction;

[0020] The weight matrix preprocessing unit is further used to perform a corresponding matrix transposition on the weight matrix data currently cached in the weight matrix cache unit based on the preset matrix multiplication instruction;

[0021] Wherein, the preset matrix multiplication instruction includes a first matrix multiplication instruction, a second matrix multiplication instruction, a third matrix multiplication instruction and a fourth matrix multiplication instruction;

[0022] The first matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which both the activation matrix data and the weight matrix data are not transposed;

[0023] The second matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which the activation matrix data is not transposed and the weight matrix data is transposed;

[0024] The third matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which the activation matrix data is transposed and the weight matrix data is not transposed;

[0025] The fourth matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which both the activation matrix data and the weight matrix data are transposed.

[0026] According to a tensor core system for a large language model provided by the present invention, the weight matrix processing component also includes a dequantization unit, and the dequantization unit is used to dequantize the target weight matrix when it is determined that the weight matrix data is in UINT4 number system or UINT8 number system, so as to obtain the target weight matrix after dequantization.

[0027] According to a tensor core system for a large language model provided by the present invention, the tensor core computing component is further used for:

[0028] A matrix floating-point multiplication and addition operation is performed on the target activation matrix and the target weight matrix after the inverse quantization process to obtain a floating-point multiplication and addition operation result having the inverse quantization process result.

[0029] According to a tensor core system for a large language model provided by the present invention, the tensor core computing component is further used for:

[0030] Based on the preset floating-point number system, the accumulation matrix stored in the memory and the floating-point multiplication and addition operation result with the inverse quantization processing result are accumulated to obtain the target result matrix.

[0031] According to a tensor core system for a large language model provided by the present invention, the accumulation matrix and result processing component includes an accumulation matrix AXI bus control unit, an accumulation matrix buffer unit, a result matrix output floating point conversion unit and a result matrix output buffer unit, wherein:

[0032] The accumulation matrix AXI bus control unit is used to send an accumulation matrix read request to the memory to obtain the accumulation matrix stored in the memory; and after sending a data write request to the memory, write the result matrix to be written in the result matrix output buffer unit to the memory;

[0033] The accumulation matrix buffer unit is used to cache the accumulation matrix;

[0034] The result matrix output floating point number conversion unit is used to perform floating point number system conversion processing on the target result matrix based on a preset floating point number system output format to obtain the result matrix to be written;

[0035] The result matrix output buffer unit is used to buffer the result matrix to be written.

[0036] According to a tensor core system for a large language model provided by the present invention, when the floating-point number system of the activation matrix data is the FP16 number system or the BF16 number system, and the floating-point number system of the weight matrix data is the quantized UINT4 number system or the UINT8 number system, the preset floating-point number system output format corresponding to the result matrix to be written is the FP16 number system or the BF16 number system;

[0037] When the floating point number system of the activation matrix data and the weight matrix data is FP16 number system or BF16 number system, the preset floating point number system output format corresponding to the result matrix to be written is FP16 number system or BF16 number system;

[0038] When the floating point number system of the activation matrix data and the weight matrix data is FP8, the preset floating point number system output format corresponding to the result matrix to be written is FP16 or BF16;

[0039] When the floating point number system of the activation matrix data and the weight matrix data is FP6 number system or FP4 number system, the preset floating point number system output format corresponding to the result matrix to be written is FP16 number system or BF16 number system.

[0040] The present invention also provides a hardware chip, comprising the above-mentioned tensor core system for a large language model.

[0041] The tensor core system and hardware chip for a large language model provided by the present invention construct an activation matrix processing component and a weight matrix processing component that can perform matrix transposition. When executing a matrix multiplication task, the tensor core independently completes data transposition preprocessing and subsequent multiplication and addition operations, saving the bandwidth required for moving the transposed data, reducing the chip area and power consumption, and making the data flow for transposition, multiplication and addition in the tensor core smoother and more efficient, which is more conducive to improving the chip computing power utilization. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0043] Figure 1 A structural diagram of a tensor core system for a large language model provided by the present invention;

[0044] Figure 2 This is a schematic diagram of the structure of the hardware chip provided in this application. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] When the Graphics Processing Unit (GPU) performs large language model inference calculations, common complex operations, such as quantization and dequantization of weight data, and matrix multiplication after data transposition, are all completed by vector cores or by vector cores and tensor cores together. However, due to the universal design of the GPU, the ratio of its vector core computing power to tensor core computing power does not match the ratio of vector core computing power to tensor core computing power required for large language model inference calculations. This is mainly reflected in that when the GPU is performing quantization and dequantization and matrix data transposition, if the vector core computing power is insufficient, the tensor core can only be in a waiting state, causing the huge computing power of the tensor core to be idle. In addition, when using the GPU for large language model inference calculations, the matrix data must be pre-processed by the vector core before entering the tensor core to start matrix multiplication operations, resulting in the data transfer between the vector core, tensor core and static random access memory (SRAM) occupying the SRAM bandwidth during the whole process, increasing the chip power consumption.

[0047] Ascend 910 is aimed at AI application fields and is mainly suitable for application scenarios such as AI training and autonomous driving. The main computing power of Ascend 910 is provided by a Cube Core that supports 16×16×16 (in FP16 number system) and 16×32×16 (in INT8 number system). However, Ascend 910 also has problems similar to GPUs, that is, Cube Core can only perform matrix multiplication operations, and vector core support is required for transposition and quantization and dequantization operations. Moreover, Ascend 910 is not a dedicated chip designed for large language models. There is a mismatch between the computing power of its vector core and Cube Core when performing large language model inference calculations, resulting in low chip computing power utilization, bandwidth and power consumption waste.

[0048] In addition, in large language models, activation matrix data is usually stored in FP16 or BF16 format, and weight matrix data is usually stored in UINT4 or UINT8 format. Ascend 910's Cube Core does not support the BF16 format commonly used in large language models, and the supported INT8 matrix operations do not match those in large language models, resulting in a waste of chip power consumption and area.

[0049] The tensor cores inside the Tensor Processing Unit (TPU) chip use a systolic array to arrange data flow. The advantage of the systolic array is that the data flow is simple and the structure is simple, so the layout and wiring of the front-end design and back-end of the TPU are relatively simple and easy to do. In addition, all the operations in the systolic array group are multiplication and addition operations, and one data is multiplied and added each time, so the calculation accuracy is high when the calculation is performed in full accordance with the IEEE 754 specification. However, the data needs to be moved from one end of the systolic array to the other end in sequence to complete the calculation, resulting in a large delay. In addition, the data in the systolic array is added sequentially, and the data path cannot be efficiently used to simplify the calculation, that is, parallel calculation cannot be performed, resulting in the TPU's computing power area ratio being significantly lower than the 3D tensor core of the GPU. Based on the current large language model inference calculation requires high computing power and is insensitive to errors, the systolic array tensor core of the TPU mode does not meet the current large language model requirements.

[0050] Based on the problems existing in the above-mentioned prior art, the large language model architecture based on the Transformer Decoder-Only technical route has become mature, and a dedicated inference chip for this architecture is urgently needed to serve the reasoning scenario of the large language model with higher efficiency and lower cost. Therefore, based on the Transformer Decoder-Only large language model architecture, the present invention constructs a fully matching dedicated inference chip tensor core architecture system, and the customized design supports mixed number system reasoning calculation of large language models, thereby solving the problem of mismatch between the computing power ratio of vector cores and tensor cores in large language model application scenarios, improving chip computing power, computing power utilization and bandwidth, and reducing chip power consumption.

[0051] Figure 1 The structure diagram of the tensor core system for a large language model provided by the present invention is as follows: Figure 1 As shown, the present invention provides a tensor core system for a large language model, comprising:

[0052] The activation matrix processing component 101 is used to cache the activation matrix data read from the memory, and perform activation matrix preprocessing on the cached activation matrix data to obtain a target activation matrix.

[0053] In the present invention, the activation matrix processing component 101 is a hardware component specially designed for processing the activation matrix. The activation matrix processing component 101 reads the data of the activation matrix from the main memory of the system (such as SRAM). These data may be the results of calculations of the previous neural network layer and are stored in the memory for use by subsequent layers.

[0054] In the present invention, in order to reduce the delay of accessing the main memory and improve the data access speed, a corresponding buffer area can be set in the activation matrix processing component 101 to temporarily store the read activation matrix data. Further, the activation matrix data after the buffer is preprocessed, and the activation matrix preprocessing step includes operations such as data transposition and decompression, and then the preprocessed target activation matrix is ​​output through the activation matrix processing component 101 to match the input requirements of the subsequent processing steps.

[0055] The weight matrix processing component 102 is used to cache the weight matrix data read from the memory, and perform weight matrix preprocessing on the cached weight matrix data to obtain a target weight matrix.

[0056] In the present invention, the weight matrix processing unit 102 is a hardware component specially designed for processing weight matrices. In a neural network, a weight matrix is ​​a set of parameters that connect neurons in different layers, and these parameters are adjusted during the training process to minimize errors. The weight matrix processing unit 102 reads the weight matrix data from the main memory (such as SRAM) of the system. These weight matrix data are important parameters that the neural network needs to access during training or reasoning.

[0057] In the present invention, in order to reduce the delay in accessing the main memory and improve the data access speed, a corresponding cache area can be set in the weight matrix processing component 102 to temporarily store the read weight matrix data. Furthermore, the cached weight matrix data is subjected to weight matrix preprocessing, and the weight matrix preprocessing step includes operations such as data transposition and decompression, and then the preprocessed target weight matrix is ​​output through the weight matrix processing component 102 to ensure that the weight matrix meets the requirements of subsequent calculations. Preferably, in the present invention, for weight matrix data of some specific number systems (such as UINT4 and UINT8), the weight matrix processing component 102 is also provided with an inverse quantization unit, which can realize the corresponding inverse quantization operation on the weight matrix data of these specific number systems.

[0058] The tensor core operation component 103 is used to perform matrix floating-point multiplication and addition operations on the target activation matrix and the target weight matrix to obtain floating-point multiplication and addition results, and based on a preset floating-point number system, accumulate the accumulation matrix stored in the memory and the floating-point multiplication and addition results to obtain a target result matrix.

[0059] In the present invention, the tensor core computing unit 103 is a hardware component specially designed to perform large-scale matrix and tensor operations. The tensor core can efficiently perform operations such as matrix multiplication and addition, especially for large-scale tensor operations in deep learning, and has a significant acceleration effect.

[0060] In the present invention, the tensor core operation component 103 first performs matrix floating-point multiplication and addition operations on the target activation matrix and the target weight matrix to obtain floating-point multiplication and addition results. The accumulation matrix is ​​a matrix that stores the previous calculation results. The tensor core operation component 103 performs accumulation operations on the results of the floating-point multiplication and addition operations and the corresponding elements in the accumulation matrix based on the preset floating-point number system, and then outputs the accumulated target result matrix. The preset floating-point number system represents the precision and format used when performing floating-point operations. In the present invention, the preset floating-point number system can be the FP32 number system. The present invention selects a suitable floating-point number system to reduce the computational complexity and memory usage while ensuring the calculation accuracy, thereby improving the overall performance.

[0061] The present invention is based on the Transformer Decoder-Only large language model architecture, and is a tensor core architecture customized for a large language model AI reasoning chip, that is, the chip application scenario is clear (for reasoning of large language models), and the tensor core function design is targeted (supporting the input and output of mixed floating-point number matrices of large language models and high-performance parallel computing). The AI ​​chips in the existing related technologies are either general AI chips or special AI chips that do not fully match the reasoning of large language models. Matrix transposition is performed in the vector core outside the tensor core. After the transposition is completed, the matrix data enters the tensor core for multiplication and addition operations. Due to the mismatch between the vector core and the tensor core computing power, there are problems in the existing related technologies that cause waste of tensor core computing power, waste of SRAM bandwidth, and increase chip power consumption. In the present invention, the tensor core includes a preprocessing unit that can perform matrix transposition (such as the preprocessing unit of the activation matrix and the weight matrix). When performing matrix multiplication tasks, the tensor core independently completes the data transposition preprocessing and subsequent multiplication and addition operations. Compared with related technical solutions, the present invention saves the bandwidth required for transposing data and reduces chip area and power consumption. In addition, the data flow of "transposition-multiplication and addition" in the tensor core is smoother and more efficient, which is more conducive to improving the utilization of chip computing power.

[0062] The accumulation matrix and result processing component 104 is used to perform floating point number conversion processing on the target result matrix to obtain a result matrix to be written, and write the result matrix to be written into the memory.

[0063] In the present invention, the accumulation matrix and result processing component 104 performs corresponding floating point number system conversion processing on the target result matrix based on the preset output data number system, and then writes the obtained result matrix to be written into the memory for subsequent matrix calculation process.

[0064] The tensor core system for a large language model provided by the present invention constructs an activation matrix processing component and a weight matrix processing component that can perform matrix transposition. When executing a matrix multiplication task, the tensor core independently completes data transposition preprocessing and subsequent multiplication and addition operations, saving the bandwidth required for moving the transposed data, reducing the chip area and power consumption, and making the data flow for transposition, multiplication and addition in the tensor core smoother and more efficient, which is more conducive to improving the chip computing power utilization.

[0065] On the basis of the above embodiment, the activation matrix processing component includes an activation matrix AXI bus control unit, an activation matrix buffer unit and an activation matrix pre-processing unit, wherein:

[0066] The activation matrix AXI bus control unit is used to determine the activation matrix storage address information corresponding to the activation matrix data in the memory according to the dimension information of the activation matrix data, and send an activation matrix read request to the memory, wherein the activation matrix read request is constructed according to the activation matrix storage address information;

[0067] The activation matrix buffer unit is used to cache the activation matrix data;

[0068] The activation matrix preprocessing unit is used to perform matrix transposition on the activation matrix data currently cached in the activation matrix cache unit, and based on the floating-point number system of the activation matrix data currently cached in the activation matrix cache unit, perform corresponding decompression processing on the floating-point elements in the activation matrix data after the matrix transposition to obtain the target activation matrix.

[0069] In the present invention, the activation matrix processing component is used to control and preprocess the input data corresponding to the activation matrix. Specifically, the activation matrix processing component includes an activation matrix AXI bus control unit, an activation matrix buffer unit and an activation matrix preprocessing unit, wherein the activation matrix AXI bus control unit first calculates the activation matrix data address (i.e., the activation matrix storage address information) according to the control instruction (such as triggering the activation matrix processing component to start executing the relevant control and preprocessing of the activation matrix data), and then sends a read data request (i.e., an activation matrix read request) to the SRAM according to the activation matrix data address, and reads out the activation matrix data stored in the SRAM. At this time, after determining that the activation matrix data is ready, the activation matrix processing component sends a ready signal to the tensor core operation component. In the present invention, the activation matrix AXI bus control unit is responsible for managing the communication with the external memory (such as SRAM), and the activation matrix AXI bus control unit calculates the storage address of the activation matrix data in the SRAM according to the dimensions of the activation matrix (such as the number of rows and columns) and the storage layout of the SRAM (such as the address mapping rule). After calculating the storage address, the activation matrix AXI bus control unit generates a read request according to the AXI bus protocol and sends it through the SRAM datapath. After the SRAM responds to the read request, it transmits the required activation matrix data back through the datapath, and the data will be stored in the activation matrix buffer.

[0070] In the present invention, the functions of the activation matrix preprocessing unit include transposition of activation matrix data and unified decompression of mixed number system floating point numbers. The activation matrix buffer unit is a storage area for temporarily storing activation matrix data required for the calculation process (these activation matrix data are read from SRAM). The present invention can increase data access speed and reduce memory access delay by activating the matrix buffer unit, thereby accelerating the calculation process.

[0071] Furthermore, the present invention performs matrix transposition on the activation matrix data currently cached to the activation matrix cache unit. Matrix transposition refers to swapping the rows and columns of the matrix, that is, the i-th row and j-th column element of the matrix becomes the j-th row and i-th column element after transposition. In this process, the activation matrix preprocessing unit performs a transposition operation on the activation matrix data in the cache, thereby providing a transposed matrix format for subsequent matrix operations.

[0072] In the field of deep learning, in order to save storage space and increase the range of data representation, data is usually represented in floating point form. When performing calculations, it is necessary to decompress the floating point numbers and restore them to a format suitable for calculation. The present invention decompresses the floating point elements in the transposed matrix according to the floating point number system (such as 16-bit floating point numbers, etc.) of the activation matrix data in the cache, so as to restore the compressed data to its original format or precision for subsequent calculations and analysis. After the matrix transposition and decompression processing, the target activation matrix is ​​finally obtained, which is used in the subsequent calculation process of the tensor core component.

[0073] In the present invention, the activation matrix preprocessing unit performs decompression processing on the floating-point elements in the transposed matrix. In the matrix calculation process of different batches, these floating-point numbers may be stored in different number systems (such as 16 bits, etc.) and compression formats. The activation matrix preprocessing unit will perform corresponding decompression processing on the floating-point elements in the activation matrix data after the matrix transposition according to the floating-point number system of the current activation matrix data, thereby obtaining the target activation matrix of the required accuracy and format.

[0074] On the basis of the above embodiment, the weight matrix processing component includes a weight matrix AXI bus control unit, a weight matrix buffer unit and a weight matrix preprocessing unit, wherein:

[0075] The weight matrix AXI bus control unit is used to determine the weight matrix storage address information corresponding to the weight matrix data in the memory according to the dimension information and quantization information of the weight matrix data, and send a weight matrix read request to the memory, wherein the weight matrix read request is constructed according to the weight matrix storage address information;

[0076] The weight matrix buffer unit is used to cache the weight matrix data;

[0077] The weight matrix preprocessing unit is used to perform matrix transposition on the weight matrix data currently cached in the weight matrix cache area unit, and based on the floating-point number system of the weight matrix data currently cached in the weight matrix cache area unit, perform corresponding decompression processing on the floating-point elements in the weight matrix data after the matrix transposition to obtain the target weight matrix.

[0078] In the present invention, the weight matrix processing component is used to control and preprocess the input data corresponding to the weight matrix. Specifically, the weight matrix processing component includes a weight matrix AXI bus control unit, a weight matrix buffer unit and a weight matrix preprocessing unit, wherein the weight matrix AXI bus control unit first calculates the weight matrix data address (i.e., the weight matrix storage address information) according to the control instruction (such as triggering the weight matrix processing component to start executing the relevant control and preprocessing of the weight matrix data), and then sends a read data request (i.e., a weight matrix read request) to the SRAM according to the weight matrix data address, and reads out the weight matrix data stored in the SRAM. At this time, after determining that the weight matrix data is ready, the weight matrix processing component sends a ready signal to the tensor core operation component. In the present invention, the weight matrix AXI bus control unit is responsible for managing communication with an external memory (such as SRAM), and the weight matrix AXI bus control unit calculates the storage address of the weight matrix data in the SRAM according to the dimensions of the weight matrix (such as the number of rows and columns), quantization information, and the storage layout of the SRAM (such as address mapping rules). After calculating the storage address, the weight matrix AXI bus control unit generates a read request according to the AXI bus protocol and sends it through the SRAM datapath. After the SRAM responds to the read request, it transmits the required weight matrix data back through the datapath, and these data will be stored in the weight matrix cache.

[0079] In the present invention, the functions of the weight matrix preprocessing unit include transposition of weight matrix data and unified decompression of mixed number system floating point numbers. The weight matrix cache unit is a storage area for temporarily storing weight matrix data required for the calculation process (these weight matrix data are read from SRAM). The present invention can improve data access speed and reduce memory access delay through the weight matrix cache unit, thereby accelerating the calculation process.

[0080] Furthermore, the present invention performs matrix transposition on the weight matrix data currently cached in the weight matrix cache unit. In this process, the weight matrix preprocessing unit performs a transposition operation on the weight matrix data in the cache, thereby providing a transposed matrix format for subsequent matrix operations.

[0081] In the field of deep learning, in order to save storage space and increase the range of data representation, floating point numbers are usually used to represent data. When performing calculations, floating point numbers need to be decompressed and restored to a format suitable for calculation. The present invention decompresses the floating point elements in the transposed matrix according to the floating point number system (such as 16-bit floating point numbers, etc.) of the weight matrix data in the cache, so as to restore the compressed data to its original format or precision for subsequent calculations and analysis. It should be noted that when the weight matrix is ​​in the UINT number system, the weight matrix preprocessing unit decompresses the quantized scaling factor and zeroing factor, and the specific numbers of UINT are not decompressed, but the specific numbers of UINT are dequantized by the dequantization unit. After the matrix transposition and decompression processing, the target weight matrix is ​​finally obtained, and this target weight matrix is ​​used for the calculation process of the subsequent tensor core component.

[0082] In the present invention, the weight matrix preprocessing unit decompresses the floating-point elements in the transposed matrix. In the matrix calculation process of different batches, these floating-point numbers may be stored in different number systems (such as 16 bits, etc.) and compression formats. The weight matrix preprocessing unit will perform corresponding decompression processing on the floating-point elements in the weight matrix data after the matrix transposition according to the floating-point number system of the current weight matrix data, thereby obtaining the target weight matrix of the required accuracy and format.

[0083] On the basis of the above embodiment, the activation matrix preprocessing unit is further used to perform corresponding matrix transposition on the activation matrix data currently cached to the activation matrix cache unit based on a preset matrix multiplication instruction;

[0084] The weight matrix preprocessing unit is further used to perform a corresponding matrix transposition on the weight matrix data currently cached in the weight matrix cache unit based on the preset matrix multiplication instruction;

[0085] Wherein, the preset matrix multiplication instruction includes a first matrix multiplication instruction, a second matrix multiplication instruction, a third matrix multiplication instruction and a fourth matrix multiplication instruction;

[0086] The first matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which both the activation matrix data and the weight matrix data are not transposed;

[0087] The second matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which the activation matrix data is not transposed and the weight matrix data is transposed;

[0088] The third matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which the activation matrix data is transposed and the weight matrix data is not transposed;

[0089] The fourth matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which both the activation matrix data and the weight matrix data are transposed.

[0090] In the present invention, the activation matrix preprocessing unit performs a corresponding matrix transposition operation on the activation matrix data currently cached in the activation matrix buffer unit. Specifically, the activation matrix preprocessing unit determines whether it is necessary to transpose the activation matrix data according to a preset matrix multiplication instruction. For example, in some neural network layers, the activation matrix data may need to be transposed in order to correctly align with the weight matrix data for multiplication operations. By performing these transposition operations, the activation matrix preprocessing unit ensures the correctness and efficiency of subsequent matrix multiplication operations.

[0091] The weight matrix preprocessing unit is similar to the activation matrix preprocessing unit. It also performs corresponding transposition operations on the weight matrix data currently cached in the weight matrix cache unit based on preset matrix multiplication instructions.

[0092] In the present invention, the preset matrix multiplication instructions are a set of instructions for guiding the activation matrix preprocessing unit, the weight matrix preprocessing unit and the tensor core operation unit to operate, and these instructions include four different matrix multiplication instructions, each of which corresponds to a different combination of matrix transposition and multiplication operations:

[0093] The first matrix multiplication instruction (Gemm), after receiving the activation matrix preprocessing unit and the weight matrix preprocessing unit, will not perform matrix transposition, so that the tensor core operation component will perform a matrix multiplication process without transposing the activation matrix data and the weight matrix data according to this instruction. This instruction belongs to a conventional matrix multiplication operation and is suitable for most standard neural network layers.

[0094] The second matrix multiplication instruction (Gemt), after the activation matrix preprocessing unit and the weight matrix preprocessing unit receive this instruction, only the weight matrix preprocessing unit performs the transposition of the weight matrix data, so that the tensor core operation component performs the matrix multiplication process according to this instruction without transposing the activation matrix data and transposing the weight matrix data.

[0095] The third matrix multiplication instruction (Getm), after receiving the activation matrix preprocessing unit and the weight matrix preprocessing unit, only the activation matrix preprocessing unit performs the transposition of the activation matrix data, so that the tensor core operation component performs the matrix multiplication process of transposing the activation matrix data according to this instruction, and the weight matrix data is not transposed.

[0096] The fourth matrix multiplication instruction (Gett), after the activation matrix preprocessing unit and the weight matrix preprocessing unit receive this instruction, the weight matrix preprocessing unit and the tensor core operation component both perform the corresponding matrix transposition, thereby causing the tensor core operation component to perform a matrix multiplication process in which both the activation matrix data and the weight matrix data are transposed according to this instruction.

[0097] Based on the above embodiment, the weight matrix processing component also includes a dequantization unit, which is used to dequantize the target weight matrix when it is determined that the weight matrix data is in UINT4 number system or UINT8 number system, so as to obtain the target weight matrix after dequantization.

[0098] In the present invention, the dequantization unit in the weight matrix processing component is used to restore (or approximately restore) the quantized weight matrix data to its original numerical range or precision. Quantization is a technique to reduce the number of bits required for data representation, which is often used to compress data size, reduce storage requirements and speed up the calculation process, but it will introduce a certain loss of precision. Dequantization is the inverse operation of this process, which aims to restore the data to a state as close to its original precision as possible.

[0099] Furthermore, when it is determined that the target weight matrix is ​​in UINT4 or UINT8 format, it means that the storage format of the matrix data needs to be dequantized, and the dequantization unit performs the dequantization operation of the UINT4 / UINT8 data in accordance with the control register format. When it is determined that the target weight matrix is ​​in UINT4 or UINT8 format, the dequantization unit performs the corresponding dequantization operation according to the format information in the preset control register.

[0100] In the present invention, the dequantization operation involves multiplying the quantized data by a scaling factor (or dequantization factor) to restore its original numerical range; in addition, a zeroing factor is added to support the dequantization of symmetric quantization and asymmetric quantization at the same time. Among them, the scaling factor and the zeroing factor are pre-calculated according to the strategy adopted during quantization and stored in the control register or related configuration data. Through the above process, the dequantization unit can generate a target weight matrix after dequantization processing, which is closer to its original high-precision representation, thereby reducing the precision loss introduced by quantization. Based on the tensor core system, the present invention enables the dequantization unit to reuse the SRAM path and corresponding cache area of ​​the tensor core operation component, and the dequantization unit natively supports the dequantization operations of UINT4 and UINT8. When performing a dequantization operation, the dequantization result can be directly output to the tensor core operation component for subsequent calculation according to the instruction requirements.

[0101] In the present invention, the tensor core system for a large language model can not only be controlled by instructions, but also its control registers can be directly rewritten by software to achieve complex function scheduling. Specifically, the registers that can be rewritten by the present invention and the corresponding register rewriting scheme are as follows:

[0102] 1. Input activation matrix number register, which can be set to: 0 (FP16), 1 (BF16), 2 (FP8_e4m3) and 3 (FP8_e5m2);

[0103] 2. Input weight matrix number register, which can be set to: 0 (FP16), 1 (BF16), 2 (FP8_e4m3), 3 (FP8_e5m2), 4 (UINT4 row quantization), 5 (UINT4 column quantization), 6 (UINT8 row quantization) and 7 (UINT8 column quantization);

[0104] 3. Output matrix number system register (same as accumulation matrix number system), can be set to: 0 (FP16) and 1 (BF16);

[0105] 4. Initial result accumulation control register, which can be set to: 0 (no accumulation) and 1 (accumulation);

[0106] 5. Enter the activation matrix row number register and set the number of rows according to the actual situation;

[0107] 6. Enter the activation matrix column number register and set the number of columns according to the actual situation (the same as the number of rows in the input weight matrix, so just set one of them);

[0108] 7. Enter the weight matrix column number register and set the number of columns according to actual needs.

[0109] Based on the above embodiment, the tensor core computing component is further used for:

[0110] A matrix floating-point multiplication and addition operation is performed on the target activation matrix and the target weight matrix after the inverse quantization process to obtain a floating-point multiplication and addition operation result having the inverse quantization process result.

[0111] In the present invention, the tensor core operation component obtains the target activation matrix and the target weight matrix after dequantization. Then, each element of the target activation matrix is ​​multiplied by the corresponding element of the target weight matrix after dequantization, and the results of the multiplication are accumulated to obtain the final result of each output position, which is the result of multiplying the target activation matrix and the target weight matrix after dequantization. In the present invention, the tensor core operation component outputs the input data (i.e., the target activation matrix and the weight matrix) after a specific operation, wherein the weight matrix has been dequantized to restore its original numerical range or precision. Finally, the tensor core operation component outputs the floating-point multiplication and addition operation result with the dequantization processing result for subsequent calculation or processing.

[0112] Based on the above embodiment, the tensor core computing component is further used for:

[0113] Based on the preset floating-point number system, the accumulation matrix stored in the memory and the floating-point multiplication and addition operation result with the inverse quantization processing result are accumulated to obtain the target result matrix.

[0114] In the present invention, the dequantization of the weight matrix is ​​an important high-frequency requirement with a relatively fixed form in the reasoning of large language models. In the AI ​​chip design scheme of the existing related technology, the dequantization operation is usually performed by the vector core, and the operation result is transmitted to the tensor core via SRAM, and then the tensor core performs the next matrix operation. Assume that the chip area required for the vector core to process conventional tasks is M, and the additional area of ​​the vector core chip required to quickly respond to the dequantization task is m (including the calculation part and the vacant part). In the reasoning process of the existing large language model, there is no dequantization task most of the time. Therefore, the vector core chip part with an area of ​​m is idle, resulting in a waste of chip area. If the vector core does not increase the area of ​​m, and only uses the area of ​​M that matches the daily workload, when a short-term and large-scale task of dequantization of a large language model is required, due to the insufficient instantaneous bandwidth and computing power of the vector core in the existing related technology, the dequantization task is executed slowly, and the tensor core is forced to be idle, thereby affecting the processing speed of the entire chip.

[0115] In the present invention, the accumulation matrix and result processing component reads the stored accumulation matrix from the memory, and this accumulation matrix can be obtained in the previous calculation step and is used to store the intermediate result. At the same time, the accumulation matrix and result processing component obtains the floating point multiplication and addition operation result. This result is obtained by the operation of the activation matrix and the weight matrix in the above embodiment.

[0116] Furthermore, the tensor core operation component accumulates each element in the accumulation matrix with the corresponding element of the floating-point multiplication and addition operation result. After the above accumulation operation, a new matrix is ​​obtained, and each element of this matrix is ​​the result of the accumulation, that is, the target result matrix. Finally, the target result matrix is ​​stored in the memory for further use.

[0117] The present invention designs the dequantization unit into the tensor core, and the dequantization unit reuses the SRAM path and corresponding buffer area of ​​the tensor core computing component. In addition, the area of ​​the vector core is only the chip area M required for processing conventional tasks, and there is no extra chip area waste. When performing the matrix multiplication parallel operation required by the large language model, due to the large bandwidth matched by the tensor core and the support of the buffer area, the dequantization task is efficiently completed, the tensor core computing component has less waiting time and high computing power utilization. In addition, compared with the dequantization processing of the vector core, the present invention simplifies the data transfer process from the original data input from SRAM to the vector core, the dequantization processing result output from the vector core to the SRAM, and the dequantization result input from the SRAM to the tensor core in the existing related art to the original data input from the SRAM to the tensor core, and also improves the overall performance of the chip.

[0118] On the basis of the above embodiment, the accumulation matrix and result processing unit includes an accumulation matrix AXI bus control unit, an accumulation matrix buffer unit, a result matrix output floating point conversion unit and a result matrix output buffer unit, wherein:

[0119] The accumulation matrix AXI bus control unit is used to send an accumulation matrix read request to the memory to obtain the accumulation matrix stored in the memory; and after sending a data write request to the memory, write the result matrix to be written in the result matrix output buffer unit to the memory;

[0120] The accumulation matrix buffer unit is used to cache the accumulation matrix;

[0121] The result matrix output floating point number conversion unit is used to perform floating point number system conversion processing on the target result matrix based on a preset floating point number system output format to obtain the result matrix to be written;

[0122] The result matrix output buffer unit is used to buffer the result matrix to be written.

[0123] In the present invention, the accumulation matrix AXI bus control unit first calculates the storage address of the accumulation matrix in the memory according to the control instruction and the dimension of the output matrix. Then, the accumulation matrix AXI bus control unit uses the AXI bus protocol to send a read request (i.e., an accumulation matrix read request) to the data path (SRAM datapath) of the memory to obtain the accumulation matrix data stored in the memory.

[0124] After the tensor core operation unit completes the calculation, the result matrix output floating point conversion unit converts the result matrix output by the tensor core operation unit into a floating point number system according to the preset floating point number system output format. Its main functions include:

[0125] Format conversion: Convert the elements in the result matrix from one floating-point number system to another floating-point number system (such as from 32-bit floating-point numbers to 16-bit floating-point numbers) according to system requirements or user settings.

[0126] Precision Adjustment: During the conversion process, the precision of the result matrix may need to be adjusted to meet specific application requirements.

[0127] The result matrix after conversion is called the result matrix to be written, which will be sent to the result matrix output buffer unit for caching. Then, a write request is sent to the data path of the memory using the AXI bus protocol to write the result matrix data to be written cached in the result matrix output buffer unit into the memory.

[0128] After the accumulation matrix data is prepared, the AXI bus control unit sends a ready signal to the tensor core computing component, indicating that the accumulation matrix data is ready and can proceed to the next step of operation.

[0129] In the present invention, the accumulation matrix buffer unit is a temporary storage area for caching accumulation matrix data read from the memory, thereby reducing the number of accesses to the memory and improving data transmission efficiency.

[0130] On the basis of the above embodiment, when the floating point number system of the activation matrix data is FP16 number system or BF16 number system, and the floating point number system of the weight matrix data is quantized UINT4 number system or UINT8 number system, the preset floating point number system output format corresponding to the result matrix to be written is FP16 number system or BF16 number system;

[0131] When the floating point number system of the activation matrix data and the weight matrix data is FP16 number system or BF16 number system, the preset floating point number system output format corresponding to the result matrix to be written is FP16 number system or BF16 number system;

[0132] When the floating point number system of the activation matrix data and the weight matrix data is FP8, the preset floating point number system output format corresponding to the result matrix to be written is FP16 or BF16;

[0133] When the floating point number system of the activation matrix data and the weight matrix data is FP6 number system or FP4 number system, the preset floating point number system output format corresponding to the result matrix to be written is FP16 number system or BF16 number system.

[0134] In the present invention, the large language model chip built based on the tensor core system for the large language model natively supports the input and output of multiple floating-point number systems, and supports the free combination of input and output number systems. The specific combination methods are as follows:

[0135] 1) Supports input data in FP16 (sign-exponent-mantissa is 1-5-10, the same below) or BF16 (1-8-7) format. When the floating-point number format of the weight matrix data is the quantized UINT4 or UINT8 format, you can choose to output the result in FP16 (1-5-10) or BF16 (1-8-7) format.

[0136] 2) Supports input data in FP16 (1-5-10) or BF16 (1-8-7) format, and outputs results in FP16 (1-5-10) or BF16 (1-8-7) format.

[0137] 3) Supports input data in FP8_e4m3 (1-4-3) or FP8_e5m2 (1-5-2) format, performs calculations at twice the computing power of FP16 (1-5-10) format, and outputs results in FP16 (1-5-10) or BF16 (1-8-7) format.

[0138] 4) Supports input data in FP6 or FP4 format, calculations in FP8 format, and output results in FP16 (1-5-10) or BF16 (1-8-7) format.

[0139] The present invention natively supports the input data activation matrix number system of FP16 (1-5-10) or BF16 (1-8-7). The weight matrix can be quantized by row, column and block, and then obtain the UINT4 number system or UINT8 number system and the corresponding scaling factor and zeroing factor for subsequent dequantization calculation. The computing power is the same as the case where the input activation matrix and weight matrix number system are both FP16 (1-5-10).

[0140] Based on the tensor core system for large language models provided by the present invention, the constructed large language model chip supports 1GHz, and a single tensor core can provide 8T, 4T and 16TFlops computing power according to different number systems, which is specifically embodied as follows:

[0141] 1. On a 1GHz AI processor, the present invention supports the case where the input data is in the FP16 (1-5-10) number system or the BF16 (1-8-7) number system, the tensor core performs the multiplication of two matrices of 16×16 dimensions and outputs the result in a pipeline per cycle, with a computing power of 16×16×16×2 = 8KFlops / clk.

[0142] 2. On a 1GHz AI processor, the present invention natively supports the case where the input data is in the FP8_e4m3 (1-4-3) number system or the FP8_e5m2 (1-5-2) number system. The tensor core performs the multiplication of two matrices of 16×32 and 32×16 dimensions and outputs the results in a pipeline per cycle, with a computing power of 32×16×16×2 = 16KFlops / clk.

[0143] In addition, in the present invention, the large language model chip supports scheduling large matrix multiplication operations through only a single instruction, and the instruction length is variable.

[0144] Figure 2 The schematic diagram of the hardware chip structure provided for this application is as follows: Figure 2 As shown, the present invention provides a hardware chip 200, including the tensor core system 201 for a large language model described in the above embodiments.

[0145] The hardware chip provided by the present invention constructs an activation matrix processing component and a weight matrix processing component that can perform matrix transposition. When executing a matrix multiplication task, the tensor core independently completes data transposition preprocessing and subsequent multiplication and addition operations, saving the bandwidth required for transposed data movement, reducing chip area and power consumption, and making the data flow for transposition, multiplication and addition in the tensor core smoother and more efficient, which is more conducive to improving the chip computing power utilization.

[0146] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A tensor core system for a large language model, characterized in that: include: An activation matrix processing component is used to cache the activation matrix data read from the memory, and perform activation matrix preprocessing on the cached activation matrix data to obtain a target activation matrix; A weight matrix processing component is used to cache the weight matrix data read from the memory, and perform weight matrix preprocessing on the cached weight matrix data to obtain a target weight matrix; A tensor core operation component is used to perform matrix floating-point multiplication and addition operations on the target activation matrix and the target weight matrix to obtain a floating-point multiplication and addition operation result, and to accumulate the accumulation matrix stored in the memory and the floating-point multiplication and addition operation result based on a preset floating-point number system to obtain a target result matrix; The accumulation matrix and result processing component is used to perform floating point number conversion processing on the target result matrix to obtain the result matrix to be written, and write the result matrix to be written into the memory.

2. The tensor core system for a large language model according to claim 1, characterized in that The activation matrix processing component includes an activation matrix AXI bus control unit, an activation matrix buffer unit and an activation matrix preprocessing unit, wherein: The activation matrix AXI bus control unit is used to determine the activation matrix storage address information corresponding to the activation matrix data in the memory according to the dimension information of the activation matrix data, and send an activation matrix read request to the memory, wherein the activation matrix read request is constructed according to the activation matrix storage address information; The activation matrix buffer unit is used to cache the activation matrix data; The activation matrix preprocessing unit is used to perform matrix transposition on the activation matrix data currently cached in the activation matrix cache unit, and based on the floating-point number system of the activation matrix data currently cached in the activation matrix cache unit, perform corresponding decompression processing on the floating-point elements in the activation matrix data after the matrix transposition to obtain the target activation matrix.

3. The tensor core system for a large language model according to claim 2, characterized in that: The weight matrix processing component includes a weight matrix AXI bus control unit, a weight matrix buffer unit and a weight matrix preprocessing unit, wherein: The weight matrix AXI bus control unit is used to determine the weight matrix storage address information corresponding to the weight matrix data in the memory according to the dimension information and quantization information of the weight matrix data, and send a weight matrix read request to the memory, wherein the weight matrix read request is constructed according to the weight matrix storage address information; The weight matrix buffer unit is used to cache the weight matrix data; The weight matrix preprocessing unit is used to perform matrix transposition on the weight matrix data currently cached in the weight matrix cache area unit, and based on the floating-point number system of the weight matrix data currently cached in the weight matrix cache area unit, perform corresponding decompression processing on the floating-point elements in the weight matrix data after the matrix transposition to obtain the target weight matrix.

4. The tensor core system for a large language model according to claim 3, characterized in that: The activation matrix preprocessing unit is further used to perform corresponding matrix transposition on the activation matrix data currently cached to the activation matrix cache unit based on a preset matrix multiplication instruction; The weight matrix preprocessing unit is further used to perform a corresponding matrix transposition on the weight matrix data currently cached in the weight matrix cache unit based on the preset matrix multiplication instruction; Wherein, the preset matrix multiplication instruction includes a first matrix multiplication instruction, a second matrix multiplication instruction, a third matrix multiplication instruction and a fourth matrix multiplication instruction; The first matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which both the activation matrix data and the weight matrix data are not transposed; The second matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which the activation matrix data is not transposed and the weight matrix data is transposed; The third matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which the activation matrix data is transposed and the weight matrix data is not transposed; The fourth matrix multiplication instruction is used to instruct the tensor core operation component to perform a matrix multiplication process in which both the activation matrix data and the weight matrix data are transposed.

5. The tensor core system for a large language model according to claim 3, characterized in that: The weight matrix processing component also includes a dequantization unit, which is used to dequantize the target weight matrix when it is determined that the weight matrix data is in UINT4 number system or UINT8 number system, so as to obtain the target weight matrix after dequantization.

6. The tensor core system for a large language model according to claim 5, characterized in that: The tensor core computing component is also used for: The target activation matrix and the target weight matrix after the inverse quantization process are subjected to matrix floating-point multiplication and addition operations to obtain a floating-point multiplication and addition operation result having the inverse quantization process result.

7. The tensor core system for a large language model according to claim 6, characterized in that: The tensor core computing component is also used for: Based on the preset floating-point number system, the accumulation matrix stored in the memory and the floating-point multiplication and addition operation result with the inverse quantization processing result are accumulated to obtain the target result matrix.

8. The tensor core system for a large language model according to claim 1 or 7, characterized in that: The accumulation matrix and result processing unit includes an accumulation matrix AXI bus control unit, an accumulation matrix buffer unit, a result matrix output floating point number conversion unit and a result matrix output buffer unit, wherein: The accumulation matrix AXI bus control unit is used to send an accumulation matrix read request to the memory to obtain the accumulation matrix stored in the memory; and after sending a data write request to the memory, write the result matrix to be written in the result matrix output buffer unit to the memory; The accumulation matrix buffer unit is used to cache the accumulation matrix; The result matrix output floating point number conversion unit is used to perform floating point number system conversion processing on the target result matrix based on a preset floating point number system output format to obtain the result matrix to be written; The result matrix output buffer unit is used to buffer the result matrix to be written.

9. The tensor core system for a large language model according to claim 8, characterized in that: When the floating point number system of the activation matrix data is FP16 or BF16, and the floating point number system of the weight matrix data is quantized UINT4 or UINT8, the preset floating point number system output format corresponding to the result matrix to be written is FP16 or BF16; When the floating point number system of the activation matrix data and the weight matrix data is FP16 number system or BF16 number system, the preset floating point number system output format corresponding to the result matrix to be written is FP16 number system or BF16 number system; When the floating point number system of the activation matrix data and the weight matrix data is FP8, the preset floating point number system output format corresponding to the result matrix to be written is FP16 or BF16; When the floating point number system of the activation matrix data and the weight matrix data is FP6 number system or FP4 number system, the preset floating point number system output format corresponding to the result matrix to be written is FP16 number system or BF16 number system.

10. A hardware chip, characterized in that: A tensor core system for a large language model comprising any one of claims 1 to 9.

Citation Information

Patent Citations

  • Model compression method and device, electronic equipment and storage medium

    CN117371508A

  • Memory access calculation method and system for 4-bit quantization and storage medium

    CN118349778A