A tensor core operation component

By designing the decompression layer, computing layer and compression layer of tensor core computing components, the problem of uncustomized existing hardware acceleration chips is solved, efficient floating-point number operation is achieved, chip area and power consumption is saved, and the computing performance of large language models is improved.

CN119576274BActive Publication Date: 2025-07-08AEROSPACE INFORMATION RES INST CAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510143464.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-07-08
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The computing components in existing hardware acceleration chips are not customized for large language models, resulting in waste of chip power consumption and area.

Method used

A tensor core computing component is designed, including a decompression layer, a calculation layer and a compression layer. Through the decompression layer, a unified pair of 16-bit or 8-bit floating point numbers is decompressed to avoid repeated decompression of data, save chip area and power consumption, and supports inputs of FP16, BF16, FP8_e5m2 and FP8_e4m3 digital systems, and output FP16 and BF16 digital systems.

Benefits of technology

The computing power area ratio and computing power consumption ratio of the chip are improved, the data layout and software call difficulty is reduced, the computing process is optimized, and the chip cost is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119576274B_ABST
    Figure CN119576274B_ABST
Patent Text Reader

Abstract

The present invention provides a tensor core operation component, which relates to the technical field of artificial intelligence chips and includes: a decompression layer for decompressing the floating-point elements in the first matrix in a preset floating-point number system to obtain first floating-point elements in a preset bit format, or for decompressing the floating-point elements in the second matrix in a preset floating-point number system to obtain second floating-point elements in a preset bit format; a calculation layer for performing multiplication and addition operations on all the first floating-point elements corresponding to the first matrix and all the second floating-point elements corresponding to the second matrix to obtain a floating-point number in a target bit format; and a compression layer for receiving the floating-point number in the target bit format and performing compression conversion processing on the floating-point number in the target bit format to obtain a floating-point operation result in a target floating-point number system. The tensor core operation component provided by the present invention better meets the matrix multiplication calculation dimension requirements of large language models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence chips, and particularly to a tensor core arithmetic component. Background Art

[0002] In large language models, the activation matrix data is usually in FP16 number system, BF16 number system or FP8 number system. If the weight matrix data is not quantized, the number system of the weight matrix is the same as that of the activation matrix; if the weight matrix is quantized, it is usually stored in UINT8 number system and UINT4 number system.

[0003] To efficiently complete the training and inference calculations of large language models, the tensor core needs to support both the activation matrix and the weight matrix being in FP16 number system or both in BF16 number system, or the activation matrix being in FP16 number system or BF16 number system, while the weight matrix is in the above fixed-point number systems.

[0004] However, the arithmetic components in existing hardware acceleration chips are not customized for large language models, resulting in waste of chip power consumption and area. Therefore, there is an urgent need for a tensor core arithmetic component to solve the above problems. Summary of the Invention

[0005] In view of the problems existing in the prior art, the present invention provides a tensor core arithmetic component.

[0006] The present invention provides a tensor core arithmetic component, including a decompression layer, a calculation layer and a compression layer, wherein:

[0007] The decompression layer is used to decompress the floating-point elements in the first matrix in a preset floating-point number system to obtain first floating-point elements in a preset bit format, or to decompress the floating-point elements in the second matrix in the preset floating-point number system to obtain second floating-point elements in the preset bit format, where the first matrix and the second matrix are two matrices in matrix multiplication operation; the preset floating-point number system includes 16-bit floating-point number system and 8-bit floating-point number system;

[0008] The calculation layer is used to perform multiplication and addition operations on all the first floating-point elements corresponding to the first matrix and all the second floating-point elements corresponding to the second matrix to obtain a floating-point number in a target bit format;

[0009] The compression layer is used to receive the floating-point number in the target bit format and perform compression conversion processing on the floating-point number in the target bit format to obtain a floating-point operation result in a target floating-point number system.

[0010] A tensor core operation component provided according to the present invention, the decompression layer includes 512 identical decompression units; the preset bit format is a 20-bit format with bits from the 0th bit to the 19th bit, including a first preset bit format and a second preset bit format, where:

[0011] When the preset floating-point number system is the 16-bit floating-point number system, the decompression unit is used to decompress the floating-point elements in a first matrix to obtain the first floating-point elements in the first preset bit format; or, to decompress the floating-point elements in a second matrix to obtain the second floating-point elements in the first preset bit format;

[0012] Among them, the 19th bit to the 12th bit in the first preset bit format are 8-bit exponents, and the 11th bit to the 0th bit in the first preset bit format are 12-bit signed mantissas;

[0013] When the preset floating-point number system is the 8-bit floating-point number system, the decompression unit is used to decompress the floating-point elements in two first matrices to obtain the first floating-point elements in the second preset bit format; or, to decompress the floating-point elements in two second matrices to obtain the second floating-point elements in the second preset bit format;

[0014] Among them, the 19th bit to the 15th bit in the second preset bit format are 5-bit exponent bits, the 14th bit to the 10th bit in the second preset bit format are 5-bit signed mantissa bits, the 9th bit to the 5th bit in the second preset bit format are 5-bit exponent bits, and the 4th bit to the 0th bit in the second preset bit format are 5-bit signed mantissa bits.

[0015] A tensor core operation component provided according to the present invention, the calculation layer is composed of 256 identical calculation units, and the calculation unit includes a first exponent addition module, a second exponent addition module, a first maximum exponent obtaining module, a second maximum exponent obtaining module, a first mantissa right shift bit number calculation module, a second mantissa right shift bit number calculation module, a first mixed number system mantissa multiplication module, a second mixed number system mantissa multiplication module, a mantissa alignment and accumulation module, and a normalization module, where:

[0016] The first exponent addition module includes 16 groups of 8-bit adders, and the second exponent addition module includes 16 groups of 5-bit adders;

[0017] The first mantissa right shift bit number calculation module includes 16 groups of 9-bit subtractors, and the second mantissa right shift bit number calculation module includes 16 groups of 6-bit subtractors;

[0018] The first mixed - radix mantissa multiplication module includes 16 groups of 12 - bit by 12 - bit signed - number multipliers, and the second mixed - radix mantissa multiplication module includes 16 groups of 5 - bit by 5 - bit signed - number multipliers.

[0019] According to a tensor core operation component provided by the present invention, the first exponent addition module is used for:

[0020] When the preset floating - point number system is the 16 - bit floating - point number system, adding the 8 - bit exponents in the first floating - point element in the first preset bit - format and the 8 - bit exponents in the second floating - point element in the first preset bit - format, and subtracting the corresponding floating - point exponent offset value from the added result to obtain a first exponent addition calculation result;

[0021] When the preset floating - point number system is the 8 - bit floating - point number system, adding the 5 - bit exponents in the first floating - point element in the second preset bit - format and the 5 - bit exponents in the second floating - point element in the second preset bit - format, and subtracting the corresponding floating - point exponent offset value from the added result to obtain a second exponent addition calculation result;

[0022] The second exponent addition module is used for:

[0023] When the preset floating - point number system is the 16 - bit floating - point number system, performing operand isolation and register clock gating;

[0024] When the preset floating - point number system is the 8 - bit floating - point number system, adding the 5 - bit exponents in the first floating - point element in the second preset bit - format and the 5 - bit exponents in the second floating - point element in the second preset bit - format, and subtracting the corresponding floating - point exponent offset value from the added result to obtain a third exponent addition calculation result.

[0025] According to a tensor core operation component provided by the present invention, when the preset floating - point number system is the 16 - bit floating - point number system, the first maximum - exponent obtaining module outputs a first maximum exponent, where the first maximum exponent is the maximum - bit exponent in the first exponent addition calculation result output by the first exponent addition module;

[0026] When the preset floating - point number system is the 8 - bit floating - point number system, the first maximum exponent obtaining module or the second maximum exponent obtaining module outputs a second maximum exponent, where the second maximum exponent is the larger one among the first - branch maximum exponent and the second - branch maximum exponent. The first - branch maximum exponent is the maximum bit exponent in the addition calculation result of the second exponents output by the first maximum exponent obtaining module, and the second - branch maximum exponent is the maximum bit exponent in the addition calculation result of the third exponents output by the second maximum exponent obtaining module.

[0027] According to a tensor core operation component provided by the present invention, when the preset floating - point number system is the 16 - bit floating - point number system, the first mantissa right - shift bit number calculation module is used to output a first exponent difference, and the output result of the second mantissa right - shift bit number calculation module is 0, where the first exponent difference is the exponent difference between the addition calculation result of the first exponents and the first maximum exponent.

[0028] When the preset floating - point number system is the 8 - bit floating - point number system, the first mantissa right - shift bit number calculation module is used to output a second exponent difference, and the second mantissa right - shift bit number calculation module is used to output a third exponent difference, where the second exponent difference is the exponent difference between the addition calculation result of the second exponents and the second maximum exponent; the third exponent difference is the exponent difference between the addition calculation result of the third exponents and the second maximum exponent.

[0029] According to a tensor core operation component provided by the present invention, the first mixed - number - system mantissa multiplication module is used for:

[0030] When the preset floating - point number system is the 16 - bit floating - point number system, multiplying the 12 - bit signed - number mantissa in the first floating - point element with the 12 - bit signed - number mantissa in the second floating - point element in the first preset bit - format to obtain a first mantissa multiplication calculation result;

[0031] When the preset floating - point number system is the 8 - bit floating - point number system, multiplying the 5 - bit signed - number mantissa in the first floating - point element with the 5 - bit signed - number mantissa in the second floating - point element in the second preset bit - format to obtain a second mantissa multiplication calculation result;

[0032] The second mixed - number - system mantissa multiplication module is used for:

[0033] When the preset floating - point number system is the 16 - bit floating - point number system, perform operand isolation and register clock gating;

[0034] When the preset floating-point number system is the 8-bit floating-point number system, multiply the 5-bit signed mantissa in the first floating-point element of the second preset bit format by the 5-bit signed mantissa in the second floating-point element of the second preset bit format to obtain a third mantissa multiplication result.

[0035] According to a tensor core operation component provided by the present invention, the mantissa alignment and accumulation module is used for:

[0036] When the preset floating-point number system is the 16-bit floating-point number system, according to the first exponent difference, arithmetically right-shift the mantissa bits in multiple groups of the first mantissa multiplication results, align the multiple groups of arithmetically right-shifted first mantissa multiplication results from the high bit, and accumulate the alignment results to obtain a first mantissa alignment and accumulation result;

[0037] When the preset floating-point number system is the 8-bit floating-point number system, according to the second exponent difference, arithmetically right-shift the mantissa bits in multiple groups of the second mantissa multiplication results; according to the third exponent difference, arithmetically right-shift the mantissa bits in multiple groups of the third mantissa multiplication results; and align the multiple groups of arithmetically right-shifted second mantissa multiplication results and the multiple groups of arithmetically right-shifted third mantissa multiplication results from the high bit, and accumulate the alignment results to obtain a second mantissa alignment and accumulation result.

[0038] According to a tensor core operation component provided by the present invention, the normalization module is used for:

[0039] Determine a target sign bit value according to the highest sign bit in the first mantissa alignment and accumulation result or the second mantissa alignment and accumulation result;

[0040] Convert the mantissa in the first mantissa alignment and accumulation result or the second mantissa alignment and accumulation result into an unsigned number, and obtain an index value corresponding to the position of the first 1 in the unsigned number in the direction from the highest bit to the lowest bit;

[0041] Based on the index value, shift the unsigned number to the left to obtain a target mantissa bit;

[0042] Construct a target exponent bit according to the first maximum exponent and the index value; or construct the target exponent bit according to the second maximum exponent and the index value;

[0043] Obtain the target bit format floating-point number according to the target sign bit value, the target exponent bit, and the target mantissa bit;

[0044] Wherein, the target bit format floating-point number is a 20-bit floating-point number.

[0045] A tensor core operation component provided by the present invention, the compression layer includes 256 identical floating-point compression units, and the floating-point compression units are used for:

[0046] When the target floating-point number system is the FP16 number system, based on the IEEE754 standard, compress the floating-point number in the target bit format into the floating-point operation result in the FP16 number system;

[0047] When the target floating-point number system is the BF16 number system, based on the IEEE754 standard, compress the floating-point number in the target bit format into the floating-point operation result in the BF16 number system;

[0048] Wherein, during the compression process, if the target exponent bit in the target bit format floating-point number is greater than or equal to the preset floating-point bit, the floating-point compression unit outputs an infinite number; if the target exponent bit in the target bit format floating-point number is less than or equal to 0, the floating-point compression unit outputs a subnormal number; if the target bit format floating-point number is a NaN, the floating-point compression unit outputs a quiet NaN, wherein the preset floating-point bit is the maximum value corresponding to the exponent bit of the target floating-point number system.

[0049] The tensor core operation component provided by the present invention constructs a tensor core operation component architecture with a decompression layer, a calculation layer and a compression layer. The decompression layer uniformly decompresses 16-bit floating-point numbers or 8-bit floating-point numbers, realizing that the decompression layer is specialized in data decompression, avoiding resource waste caused by repeated data decompression, and saving chip area and power consumption. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0051] Figure 1 It is a schematic structural diagram of the tensor core operation component provided by the present invention;

[0052] Figure 2 It is a schematic architecture diagram of the calculation unit provided by the present invention;

[0053] Figure 3 It is a schematic diagram of mantissa alignment and accumulation based on the mantissa alignment and accumulation module provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts fall within the scope of protection of the present invention.

[0055] As a general-purpose artificial intelligence chip, the Graphics Processing Unit (GPU) has an architecture that supports general architectures for image processing, scientific computing, artificial intelligence model training, and inference. The existing GPUs mainly include three series: V100, A100, and H100. However, the three dimensions of matrix multiply-accumulate supported by the tensor cores of these series of GPUs are not symmetric and all contain the dimension "4". Therefore, the existing GPU tensor cores are relatively "flat". Perhaps to make the GPU more general and considering the scenario of tensor cores applied to image operations, the AI computing card follows the dimension design of the graphics card that includes the dimension "4". In addition, since the GPU is a general-purpose artificial intelligence chip, its architecture design needs to take into account supporting multiple applications. Therefore, to avoid waste of GPU performance caused by a small matrix multiplication dimension in some application scenarios while the tensor core supports a large matrix multiply-accumulate dimension, the matrix multiply-accumulate dimension supported by the tensor cores of existing GPUs cannot be made very large.

[0056] During the training and inference processes of large language models, except for some attention algorithms, more than 90% of the chip computing power is required to support matrix multiplication of large dimensions (dimensions exceeding 1024). Therefore, when the GPU is used for the calculation of large language models, the tensor core calculation with such a small dimension as "4" will cause problems such as complex data arrangement and inconvenient software calls. It is known that the larger the tensor core accumulation dimension, the more simplified operations can be carried out in the floating-point operation intermediate accumulation process. Therefore, the computing power area ratio and the computing power power consumption ratio are also larger, and the tensor core performance is better. For a general-purpose artificial intelligence chip such as the GPU, its architecture design needs to take into account supporting multiple applications, resulting in a relatively small tensor core dimension in the existing design of the GPU, limited simplified operations in the accumulation process, and low computing power area ratio and computing power power consumption ratio.

[0057] Ascend 910 is targeted at the field of AI applications and is mainly applicable to application scenarios such as AI training and autonomous driving. The main computing power of Ascend 910 is provided by a Cube Core that supports 16×16×16 (in FP16 number system) and 16×32×16 (in INT8 number system). The Cube Core does not support the BF16 number system commonly used in large language models, and the activation matrix and weight matrix supported by the Cube Core are both matrix operation methods in INT8 number system, which are also not compatible with the operations in large language models.

[0058] The tensor cores inside the Tensor Processing Unit (TPU) chips arrange the data flow in the way of a systolic array. The advantage of the systolic array scheme lies in the simplicity of the data flow and the structure, so the front-end design and the layout and routing of the back-end of the TPU are relatively easy to do. And all operations within the systolic array are multiply-accumulate operations, with one data being multiplied and accumulated each time, so the operations can be fully carried out in accordance with the IEEE 754 standard, with relatively high operation precision. However, in the systolic array scheme, the data needs to move sequentially from one end of the systolic array to the other end to complete the operations, resulting in a relatively large latency. Moreover, the data in the systolic array is added sequentially, and it is impossible to efficiently utilize the data path to simplify the calculations, that is, parallel calculations cannot be performed. In addition, the existing TPUs have poor compatibility with some mainstream model frameworks in the field of artificial intelligence, and developers need additional costs to adapt to the architecture of the TPUs when using them.

[0059] Although the existing TPU systolic array scheme has high precision, it has a relatively large latency, and the data in the systolic array is added sequentially, with the accumulation dimension being only 1. Therefore, the computing power area ratio and the computing power power consumption ratio are very low. Under the current requirements of high computing power for large language model calculations and the insensitivity to errors in some scenarios, the TPU is not applicable. In addition, the existing mainstream large model algorithms and firmware are all developed from GPUs. Using the systolic array tensor cores of TPUs, which are completely different from the tensor cores of GPUs, will lead to complex software migration, complex firmware development, and high learning costs.

[0060] In view of the problems existing in the above-mentioned prior art, the large language model architecture based on the Transformer Decoder-Only technical route has become mature, and there is an urgent need for dedicated chips for this architecture to serve the training and inference of large language models with higher efficiency and lower cost. The tensor core operation component customized in the present invention for large language models supports number systems that are exactly matched with the number systems required by large language models. The supported input floating-point number systems include FP16 (sign-exponent-mantissa is 1-5-10) number system that complies with IEEE 754, BF16 (1-7-8) number system, FP8_e5m2 (1-5-2) number system, and FP8_e4m3 (1-4-3) number system that does not comply with IEEE 754. The supported output floating-point number systems include FP16 number system and BF16 number system. Moreover, the tensor core dimensions of 16×16×16 in the FP16 number system and 16×16×32 in the FP8 number system adopted in the present invention meet the requirements of large dimension matrix multiplication of large language models. It can not only reduce the difficulty of data arrangement and software call, but also provide a larger optimization space for the intermediate process of accumulation due to the larger accumulation dimension, improving the computing power area ratio and computing power power consumption ratio of the chip.

[0061] Compared with the existing solution of decompressing floating-point numbers first and then calculating multiplication and addition in the computing unit, in order to avoid repeated decompression of data, the present invention designs a tensor core operation component architecture of decompression layer - computing layer - compression layer, saving 15 / 16 of the area and corresponding power consumption of the decompression part. Moreover, the clear hierarchical architecture is also beneficial for backend design and development personnel to optimize timing and devices, improving the utilization rate of chip computing power. In the computing layer of the present invention, by reusing modules such as exponent addition, maximum exponent calculation, mantissa right shift bit number calculation, and mixed number system mantissa multiplication, while natively supporting mixed number system inputs of FP16 number system, BF16 number system, FP8_e5m2 number system, and FP8_e4m3 (1-4-3) number system, the chip area and power consumption are saved. In addition, in order to solve the problem of fuzzy data form in the intermediate process between the decompression layer, computing layer, and compression layer, the present invention also designs a 20-bit custom floating-point number format and corresponding processing rules. The introduction of this format improves the data processing efficiency within the tensor core operation component and improves the tensor core performance.

[0062] Figure 1 The structural schematic diagram of the tensor core operation component provided by the present invention is as Figure 1 shown. The present invention provides a tensor core operation component, including a decompression layer 101, a computing layer 102, and a compression layer 103, wherein:

[0063] The decompression layer 101 is used to decompress the floating-point elements in the first matrix of a preset floating-point number system to obtain the first floating-point elements in a preset bit format, or to decompress the floating-point elements in the second matrix of the preset floating-point number system to obtain the second floating-point elements in the preset bit format, where the first matrix and the second matrix are two matrices in a matrix multiplication operation; the preset floating-point number system includes a 16-bit floating-point number system and an 8-bit floating-point number system.

[0064] In the present invention, the supported input floating-point number systems include a 16-bit floating-point number system and an 8-bit floating-point number system, specifically: the FP16 (sign-exponent-mantissa is 1-5-10) number system, the BF16 (1-7-8) number system, the FP8_e5m2 (1-5-2) number system that conform to the IEEE 754 regulations, and the FP8_e4m3 (1-4-3) number system that does not conform to the IEEE 754 regulations.

[0065] For the tensor core matrix multiplication operation implemented in the present invention, if the matrix multiplication of a 16-bit floating-point number system (FP16 number system or BF16 number system) is implemented, its mathematical expression is: Let A, B, and C be all 16×16 floating-point matrices, C is the product of matrix A and matrix B, denoted as C = AB, where the element in the i-th row and j-th column of matrix C can be expressed as:

[0066] ;

[0067] If the matrix multiplication of an 8-bit floating-point number system (FP8_e4m3 or FP8_e5m2) is implemented, its mathematical expression is: Let A be a 16×32 floating-point matrix, B be a 32×16 floating-point matrix, and C be a 16×16 floating-point matrix. C is the product of matrix A and matrix B, denoted as C = AB, where the element in the i-th row and j-th column of matrix C can be expressed as:

[0068] ;

[0069] In the present invention, the decompression layer 101 is composed of 512 identical decompression units, which can decompress 512 16-bit floating-point numbers (i.e., the preset floating-point number system is a 16-bit floating-point number system) or 1024 8-bit floating-point numbers (i.e., the preset floating-point number system is an 8-bit floating-point number system) into a form that is easy for the floating-point calculation unit to process according to the floating-point format. Among them, the input of the decompression unit is 16 bits, and the output is 20 bits.

[0070] When the input is a 16-bit floating-point number, each decompression unit can decompress the 16-bit floating-point number elements of one matrix A (i.e., the first matrix) or matrix B (i.e., the second matrix), so as to obtain the first floating-point number elements or the second floating-point number elements in a preset bit format. When the input is an 8-bit floating-point number, each decompression unit can simultaneously decompress the 8-bit floating-point number elements of two matrix A or two matrix B, so as to obtain the first floating-point number elements or the second floating-point number elements in a preset bit format.

[0071] The computing layer 102 is configured to perform a multiply-add operation on all the first floating-point number elements corresponding to the first matrix and all the second floating-point number elements corresponding to the second matrix, so as to obtain a floating-point number in a target bit format.

[0072] In the present invention, the computing layer 102 is composed of 256 identical computing units. Each computing unit performs a multiplication operation on 16 groups of floating-point numbers respectively, and then adds the 16 groups of multiplication results to obtain a floating-point number, thereby completing a group of floating-point number multiply-add operations. When calculating 16-bit or 8-bit floating-point number multiplication, the inputs of the computing units are all 16 decompressed 20-bit custom format (i.e., the preset bit format) floating-point numbers as the elements of matrix A, and 16 decompressed 20-bit custom format floating-point numbers as the elements of matrix B. The computing unit outputs a custom-form floating-point number (i.e., the floating-point number in the target bit format) including a sign bit, an exponent bit, and a mantissa bit.

[0073] In the computing layer 102, the floating-point computing units are sequentially denoted as calc(i, j), where i, j ∈ [0, 15]. The input connection relationship between the decompression layer 101 and the computing layer 102 is:

[0074] ;

[0075] wherein, the decompression units are sequentially denoted as unpack(i, j).

[0076] The compression layer 103 is configured to receive the floating-point number in the target bit format and perform compression conversion processing on the floating-point number in the target bit format to obtain a floating-point number operation result in a target floating-point number system.

[0077] In the present invention, the compression layer 103 is composed of 256 identical floating-point compression units. Each floating-point compression unit corresponds to a floating-point computing unit. The floating-point compression unit reversely converts (i.e., performs compression conversion processing) the 20-bit custom format floating-point number output by the corresponding computing unit into a 16-bit floating-point number in the FP16 number system or the BF16 number system (i.e., the floating-point number operation result in the target floating-point number system) according to the IEEE 754 regulation.

[0078] Based on the above embodiments, the decompression layer includes 512 identical decompression units; the preset bit format is a 20-bit format with bits from the 0th bit to the 19th bit, including a first preset bit format and a second preset bit format, where:

[0079] When the preset floating-point number system is the 16-bit floating-point number system, the decompression unit is used to decompress the floating-point number elements in one of the first matrices to obtain the first floating-point number element in the first preset bit format; or, to decompress the floating-point number elements in one of the second matrices to obtain the second floating-point number element in the first preset bit format;

[0080] Among them, bits 19 to 12 in the first preset bit format are 8-bit exponents, and bits 11 to 0 in the first preset bit format are 12-bit signed mantissas;

[0081] When the preset floating-point number system is the 8-bit floating-point number system, the decompression unit is used to decompress the floating-point number elements in two of the first matrices to obtain the first floating-point number element in the second preset bit format; or, to decompress the floating-point number elements in two of the second matrices to obtain the second floating-point number element in the second preset bit format;

[0082] Among them, bits 19 to 15 in the second preset bit format are 5-bit exponent bits, bits 14 to 10 in the second preset bit format are 5-bit signed mantissa bits, bits 9 to 5 in the second preset bit format are 5-bit exponent bits, and bits 4 to 0 in the second preset bit format are 5-bit signed mantissa bits.

[0083] In the present invention, multiple decompression units are sequentially denoted as unpack(i, j), where i, j ∈ [0, 15]. When the floating-point matrix input to the tensor core operation component is a 16-bit floating-point number system in FP16 number system or BF16 number system, the decompression unit outputs a 20-bit custom format floating-point number (i.e., the first preset bit format), denoted as unpack_out[19:0], where unpack_out[19:12] (i.e., bits 19 to 12 in the first preset bit format) are 8-bit exponents, corresponding to the floating-point exponent bits of the original 16-bit floating-point number, and unpack_out[11:0] (i.e., bits 11 to 0 in the first preset bit format) is the 12-bit signed mantissa obtained by converting the original floating-point mantissa into a complement code representation according to the IEEE 754 regulation after adding the hidden bit according to the sign bit.

[0084] When the floating - point matrix of the input tensor core operation component is 8 - bit floating - point numbers in two FP8_e4m3 number systems or FP8_e5m2 number systems, the decompression unit outputs one 20 - bit custom - format floating - point number (i.e., the second preset bit - format), denoted as unpack_out[19:0]. Among them, unpack_out[19:15] (i.e., the 19th to 15th bits in the second preset bit - format) is a 5 - bit exponent, which is the exponent bit of the 8 - bit floating - point number in the original first matrix (i.e., the first of the two first matrices or the two second matrices). unpack_out[14:10] (i.e., the 14th to 10th bits in the second preset bit - format) is a 5 - bit signed - number mantissa obtained by converting the mantissa bits of the floating - point number in the original first matrix (i.e., the first of the two first matrices or the two second matrices) into two's - complement representation according to the sign - bit mark after adding the hidden bit in accordance with IEEE754 regulations. Similarly, unpack_out[9:5] (i.e., the 9th to 5th bits in the second preset bit - format) is a 5 - bit exponent, which is the exponent bit of the 8 - bit floating - point number in the original second matrix (i.e., the second of the two first matrices or the two second matrices). unpack_out[4:0] (i.e., the 4th to 0th bits in the second preset bit - format) is a 5 - bit signed - number mantissa obtained by converting the mantissa bits of the floating - point number in the original second matrix (i.e., the second of the two second matrices or the two second matrices) into two's - complement representation according to the sign - bit mark after adding the hidden bit in accordance with IEEE 754 regulations.

[0085] In the decompression layer of the present invention, 256 16 - bit floating - point numbers or 512 8 - bit floating - point numbers are uniformly decompressed. In view of the fact that the existing tensor core calculation unit needs to decompress the data before calculation, each data will be decompressed 16 times in the calculation unit, and the repeated decompression of data causes waste of resources. Therefore, the present invention constructs a decompression layer dedicated to data decompression, and a data is decompressed only once, which can save 15 / 16 of the chip area and corresponding power consumption of the decompression part. At the same time, the clear hierarchical architecture also means simpler identification of the calculation process, which is conducive to the backend design and development personnel to optimize the timing and devices.

[0086] The tensor core operation component provided by the present invention constructs a tensor core operation component architecture with a decompression layer, a calculation layer, and a compression layer. The decompression layer uniformly decompresses 16 - bit floating - point numbers or 8 - bit floating - point numbers, realizes that the decompression layer is dedicated to data decompression, avoids the waste of resources caused by repeated data decompression, and saves chip area and power consumption.

[0087] Based on the above embodiments, the computing layer is composed of 256 identical computing units. The computing unit includes a first exponent addition module, a second exponent addition module, a first maximum exponent calculation module, a second maximum exponent calculation module, a first mantissa right shift bit number calculation module, a second mantissa right shift bit number calculation module, a first mixed-radix mantissa multiplication module, a second mixed-radix mantissa multiplication module, a mantissa alignment and accumulation module, and a normalization module, where:

[0088] The first exponent addition module includes 16 groups of 8-bit adders, and the second exponent addition module includes 16 groups of 5-bit adders;

[0089] The first mantissa right shift bit number calculation module includes 16 groups of 9-bit subtractors, and the second mantissa right shift bit number calculation module includes 16 groups of 6-bit subtractors;

[0090] The first mixed-radix mantissa multiplication module includes 16 groups of 12-bit by 12-bit signed number multipliers, and the second mixed-radix mantissa multiplication module includes 16 groups of 5-bit by 5-bit signed number multipliers.

[0091] Figure 2 is a schematic diagram of the architecture of the computing unit provided by the present invention. As Figure 2 shown, in the present invention, each computing unit includes 2 exponent addition modules, 2 maximum exponent calculation modules, 2 mantissa right shift bit number calculation modules, 2 mixed-radix mantissa multiplication modules, 1 mantissa alignment and accumulation module, and 1 normalization module.

[0092] Based on the above embodiments, the first exponent addition module is used for:

[0093] When the preset floating-point number system is the 16-bit floating-point number system, add the 8-bit exponents in the first floating-point element in the first preset bit format and the 8-bit exponents in the second floating-point element in the first preset bit format, and subtract the corresponding floating-point exponent offset value from the added result to obtain a first exponent addition calculation result;

[0094] When the preset floating-point number system is the 8-bit floating-point number system, add the 5-bit exponents in the first floating-point element in the second preset bit format and the 5-bit exponents in the second floating-point element in the second preset bit format, and subtract the corresponding floating-point exponent offset value from the added result to obtain a second exponent addition calculation result;

[0095] The second exponent addition module is used for:

[0096] When the preset floating-point number system is the 16-bit floating-point number system, operand isolation and register clock gating are performed;

[0097] When the preset floating-point number system is the 8-bit floating-point number system, add the 5-bit exponents in the first floating-point element of the second preset bit format and the 5-bit exponents in the second floating-point element of the second preset bit format, and subtract the corresponding floating-point exponent offset value from the added result to obtain a third exponent addition calculation result.

[0098] In the present invention, the first exponent addition module includes 16 groups of 8-bit adders, and the second exponent addition module includes 16 groups of 5-bit adders.

[0099] When the floating-point matrix input to the tensor core operation component is in the 16-bit floating-point number system, the first exponent addition module can complete the 8-bit exponent addition of 16 groups of 16-bit floating-point numbers, and subtract the corresponding floating-point exponent offset value to obtain a first exponent addition calculation result. At this time, the second exponent addition module performs operand isolation and register clock gating to save power consumption.

[0100] When the floating-point matrix input to the tensor core operation component is 8-bit floating-point, it is necessary to reuse the 8-bit adders of the first exponent addition module to complete the 5-bit exponent addition of 16 groups of 8-bit floating-point numbers, and subtract the corresponding floating-point exponent offset value to obtain a second exponent addition calculation result. The second exponent addition module completes the 5-bit exponent addition of 16 groups of 8-bit floating-point numbers, and subtracts the corresponding floating-point exponent offset value to obtain a third exponent addition calculation result.

[0101] On the basis of the above embodiments, when the preset floating-point number system is the 16-bit floating-point number system, the first maximum exponent obtaining module outputs a first maximum exponent, where the first maximum exponent is the maximum bit exponent in the first exponent addition calculation result output by the first exponent addition module;

[0102] When the preset floating-point number system is the 8-bit floating-point number system, the first maximum exponent obtaining module or the second maximum exponent obtaining module outputs a second maximum exponent, where the second maximum exponent is the larger maximum exponent among the first branch maximum exponent and the second branch maximum exponent, the first branch maximum exponent is the maximum bit exponent in the second exponent addition calculation result output by the first maximum exponent obtaining module, and the second branch maximum exponent is the maximum bit exponent in the third exponent addition calculation result output by the second maximum exponent obtaining module.

[0103] In the present invention, when the preset floating-point number system is the 16-bit floating-point number system, the first maximum exponent obtaining module obtains the maximum value among the 16 groups of 8-bit exponent addition results calculated by the first exponent addition module, that is, the first maximum exponent.

[0104] When the preset floating-point number system is the 8-bit floating-point number system, the first maximum exponent obtaining module obtains the maximum value from the 16 groups of 8-bit / 5-bit addition results calculated by the first exponent addition module, and obtains the maximum value from the 16 groups of 5-bit exponent addition results calculated by the second exponent addition module. Then, the larger of these two maximum values (at this time, these two maximum values can be defined as the corresponding branch maximum values) is determined as the second maximum exponent, and the corresponding maximum exponent obtaining module outputs this maximum value.

[0105] Specifically, when the floating-point matrix input to the tensor core operation component is a 16-bit floating-point number, only the first maximum exponent obtaining module outputs the first maximum exponent; when the floating-point matrix input to the tensor core operation component is an 8-bit floating-point number, after the first maximum exponent obtaining module and the second maximum exponent obtaining module calculate their respective corresponding branch maximum values, by comparing the two branch maximum values, the maximum exponent obtaining module corresponding to the larger branch maximum value outputs the second maximum exponent. In the present invention, the result output by the maximum exponent obtaining module is denoted as the maximum exponent exp_max.

[0106] Based on the above embodiments, when the preset floating-point number system is the 16-bit floating-point number system, the first mantissa right shift bit number calculation module is used to output a first exponent difference, and the output result of the second mantissa right shift bit number calculation module is 0, where the first exponent difference is the exponent difference between the first exponent addition calculation result and the first maximum exponent;

[0107] When the preset floating-point number system is the 8-bit floating-point number system, the first mantissa right shift bit number calculation module is used to output a second exponent difference, and the second mantissa right shift bit number calculation module is used to output a third exponent difference, where the second exponent difference is the exponent difference between the second exponent addition calculation result and the second maximum exponent; the third exponent difference is the exponent difference between the third exponent addition calculation result and the second maximum exponent.

[0108] In the present invention, when the preset floating-point number system is the 16-bit floating-point number system, the first mantissa right shift bit number calculation module uses 16 groups of 9-bit subtractor units to output the difference between 16 groups of 9-bit exponents and the first maximum exponent. At this time, the output result of the 16 groups of 6-bit subtractor units of the second mantissa right shift bit number calculation module is 0.

[0109] When the preset floating - point number system is the 8 - bit floating - point number system, the 16 - group 9 - bit subtractors of the first mantissa right - shift bit number calculation module are multiplexed to output the difference between the 16 groups of 6 - bit exponents and the second - largest exponent; at the same time, the 16 - group 6 - bit subtractors of the second mantissa right - shift bit number calculation module are used to output the difference between the 16 groups of 6 - bit exponents and the second - largest exponent, thus a total of 32 groups of exponent differences are output.

[0110] Based on the above - mentioned embodiment, the first mixed - number - system mantissa multiplication module is used for:

[0111] When the preset floating - point number system is the 16 - bit floating - point number system, multiply the 12 - bit signed - number mantissa in the first floating - point element of the first preset bit - format and the 12 - bit signed - number mantissa in the second floating - point element of the first preset bit - format to obtain the first mantissa multiplication calculation result;

[0112] When the preset floating - point number system is the 8 - bit floating - point number system, multiply the 5 - bit signed - number mantissa in the first floating - point element of the second preset bit - format and the 5 - bit signed - number mantissa in the second floating - point element of the second preset bit - format to obtain the second mantissa multiplication calculation result;

[0113] The second mixed - number - system mantissa multiplication module is used for:

[0114] When the preset floating - point number system is the 16 - bit floating - point number system, perform operand isolation and register clock gating;

[0115] When the preset floating - point number system is the 8 - bit floating - point number system, multiply the 5 - bit signed - number mantissa in the first floating - point element of the second preset bit - format and the 5 - bit signed - number mantissa in the second floating - point element of the second preset bit - format to obtain the third mantissa multiplication calculation result.

[0116] In the present invention, the first mixed - number - system mantissa multiplication module includes 16 groups of 12 - bit × 12 - bit signed - number multipliers, outputting a 23 - bit result, and the second mixed - number - system mantissa multiplication module includes 16 groups of 5 - bit × 5 - bit signed - number multipliers, outputting a 9 - bit result.

[0117] When the preset floating - point number system is the 16 - bit floating - point number system, the present invention only uses the first mixed - number - system mantissa multiplication module to perform 16 groups of 12 - bit × 12 - bit mantissa multiplications to obtain the 23 - bit signed - number mantissa multiplication result, that is, the first mantissa multiplication calculation result. At this time, the second mixed - number - system mantissa multiplication module performs operand isolation and register clock gating to save power consumption.

[0118] When the preset floating-point number system is an 8-bit floating-point number system, 16 groups of 12-bit × 12-bit signed number multipliers of the first hybrid number system mantissa multiplication module are multiplexed at this time, and the multiplication result of 23-bit signed number mantissas is output, that is, the second mantissa multiplication calculation result; at the same time, 16 groups of 5-bit × 5-bit signed number multipliers of the second hybrid number system mantissa multiplication module are used to output the multiplication result of 9-bit signed number mantissas, that is, the third mantissa multiplication calculation result, so as to output 32 groups of mantissa multiplication results in total.

[0119] In the present invention, the exponent addition module, the maximum exponent obtaining module, and the mantissa right shift bit number calculation module inside the calculation unit only involve exponent data, and the hybrid number system mantissa multiplication module only involves mantissa data. Considering that the multiplication operation with a large bit width takes a long time, the above three exponent calculation-related modules and the mantissa multiplication module are designed to perform parallel calculations, so as to achieve full data pipelining, save registers, and reduce chip area and power consumption.

[0120] Based on the above embodiments, the mantissa alignment and accumulation module is used for:

[0121] When the preset floating-point number system is the 16-bit floating-point number system, according to the first exponent difference, the mantissa bits in multiple groups of the first mantissa multiplication calculation results are arithmetically right-shifted, and the multiple groups of first mantissa multiplication calculation results after the arithmetic right-shift of the mantissa bits are aligned starting from the high bit, and the alignment results are accumulated to obtain the first mantissa alignment and accumulation result;

[0122] When the preset floating-point number system is the 8-bit floating-point number system, according to the second exponent difference, the mantissa bits in multiple groups of the second mantissa multiplication calculation results are arithmetically right-shifted; according to the third exponent difference, the mantissa bits in multiple groups of the third mantissa multiplication calculation results are arithmetically right-shifted; and the multiple groups of second mantissa multiplication calculation results after the arithmetic right-shift of the mantissa bits and the multiple groups of third mantissa multiplication calculation results after the arithmetic right-shift of the mantissa bits are aligned starting from the high bit, and the alignment results are accumulated to obtain the second mantissa alignment and accumulation result.

[0123] In floating-point multiplication, due to different exponents, the mantissa after multiplication may need to be shifted to the right to align the exponents of the results. When the preset floating-point number system is the 16-bit floating-point number system, the mantissa alignment and accumulation module shifts the mantissa bits to the right according to the first exponent difference obtained in the above embodiments. After the mantissa bits are shifted, since the floating-point numbers may have different decimal point positions when stored, the mantissa alignment and accumulation module needs to align the mantissa parts of all floating-point numbers participating in the operation for accumulation to obtain the first mantissa alignment and accumulation result, and the output result can be denoted as mantissa[27:0].

[0124] When the preset floating - point number system is an 8 - bit floating - point number system, the mantissa alignment and accumulation module can perform a right - shift of 16 groups of 23 - bit signed numbers and a right - shift of 16 groups of 9 - bit signed numbers, and accumulate the 16 groups of 23 - bit signed numbers after the right - shift and the 16 groups of 9 - bit signed numbers after the right - shift after aligning them from the high - order bits, output a 28 - bit signed number result, obtain the second mantissa alignment and accumulation result, and record the output result as mantissa[27:0]. Figure 3 The mantissa alignment and accumulation schematic diagram provided by the present invention based on the mantissa alignment and accumulation module. The data structure for specifically implementing mantissa alignment and accumulation can be referred to Figure 3 as shown

[0125] Based on the above - mentioned embodiments, the normalization module is used for:

[0126] Determine the target sign - bit value according to the highest - order sign - bit in the first mantissa alignment and accumulation result or the second mantissa alignment and accumulation result;

[0127] Convert the mantissa in the first mantissa alignment and accumulation result or the second mantissa alignment and accumulation result into an unsigned number, and obtain the index value corresponding to the position of the first 1 in the unsigned number in the direction from the highest - order bit to the lowest - order bit;

[0128] Based on the index value, shift the unsigned number to the left to obtain the target mantissa bit;

[0129] Construct the target exponent bit according to the first maximum exponent and the index value; or, construct the target exponent bit according to the second maximum exponent and the index value;

[0130] Obtain the target bit - format floating - point number according to the target sign - bit value, the target exponent bit, and the target mantissa bit;

[0131] Wherein, the target bit - format floating - point number is a 20 - bit floating - point number in bit format.

[0132] In the present invention, the main function of the normalization module is to perform normalization processing according to the mantissa obtained by the mantissa alignment and accumulation module and the maximum exponent obtained by the maximum exponent obtaining module, so as to output a 20 - bit custom - format floating - point number, that is, the target bit - format floating - point number.

[0133] Specifically, the process of the normalization module for processing the mantissa is as follows: Obtain the highest-order sign bit of the mantissa mantissa[27:0], and the value of this sign bit is also the sign bit value of the final floating-point number output by the normalization module, which can be denoted as s (where a sign bit of 1 indicates that this floating-point number is negative, and a sign bit of 0 indicates that this floating-point number is positive); Restore the mantissa in the first mantissa alignment accumulation result or the second mantissa alignment accumulation result to an unsigned number, denoted as mantissa_out[26:0]. Further, in the direction from the highest bit to the lowest bit of mantissa_out[26:0], find the position of the first 1. Let the nth bit be the position of the first 1 (i.e., the index value); Then, shift mantissa_out[26:0] to the left by n bits, denoted as mantissa_sft[26:0]. Therefore, starting from the highest bit, the target bit format floating-point number of the finally output 20-bit bit format floating-point number is self_def[19:0], and then it is input to the floating-point compression unit for compression. Among them, the sign bit self_def

[19] =s, the exponent bit self_def[18:11]=exp_max-n+6, and the mantissa bit self_def[10:0]=mantissa_sft[26:17]. Among them, there are 7 bits before the decimal point in the 32 groups of mantissa accumulation results. To change it to the standard form with 1 bit before the decimal point, the mantissa decimal point is shifted 6 bits to the left, and the exponent should be increased by 6, that is, the number "6" in "exp_max-n+6". In the present invention, the method of obtaining the mantissa bit is not a simple truncation, but a rounding to the nearest even method.

[0134] Based on the above embodiments, the compression layer includes 256 identical floating-point compression units, and the floating-point compression unit is used for:

[0135] When the target floating-point number system is the FP16 number system, based on the IEEE754 standard, compress the target bit format floating-point number into the floating-point operation result of the FP16 number system;

[0136] When the target floating-point number system is the BF16 number system, based on the IEEE754 standard, compress the target bit format floating-point number into the floating-point operation result of the BF16 number system;

[0137] Among them, during the compression process, if the target exponent bit in the target bit format floating-point number is greater than or equal to the preset floating-point position, the floating-point compression unit outputs an infinity number; if the target exponent bit in the target bit format floating-point number is less than or equal to 0, the floating-point compression unit outputs a subnormal number; if the target bit format floating-point number is a NaN (Not a Number), the floating-point compression unit outputs a quiet NaN, where the preset floating-point position is the maximum value corresponding to the exponent bit of the target floating-point number system.

[0138] In the present invention, the 20-bit custom floating-point number to be compressed (i.e., the target bit format floating-point number) is self_def[19:0], where the sign bit is self_def

[19] =s, the exponent bit is self_def[18:11], and the unsigned mantissa is self_def[10:0]. If the final output number system is the FP16 number system, the floating-point compression result is FP16_out[15:0], FP16_out

[15] =s, FP16_out [14:10]=self_def[15:11], FP16_out [9:0]=self_def[9:0].

[0139] Specifically, during the compression process, when the final output number system is the FP16 number system, if the target exponent bit (i.e., exp_max–n+6) in the target bit format floating-point number is greater than or equal to the preset floating-point position (0x1F), that is, exp_max–n+6≥0x1F, then self_def[18:11]=0x1F, and the infinity number specified by IEEE 754 is output. The compression result is FP16_out

[15] =s, FP16_out[14:10]=5’h1F, FP16_out[9:0]=10’h0, where 0x1F is the maximum value that the exponent bit of the target floating-point number system can represent; if exp_max–n+6≤0, the subnormal (subnormal number) specified by IEEE 754 is output; if there is a NaN (Not a Number) specified by IEEE 754, the output is a quiet NaN, and the compression result is FP16_out

[15] =s, FP16_out[14:10]=5’h1F, FP16_out[9:0]=10’h201.

[0140] When the final output number system is the BF16 number system, the floating-point compression result is BF16_out[15:0], where BF16_out

[15] =s, BF16_out[14:7]=self_def[18:11], and BF16_out[6:0]=self_def[9:3]. If exp_max–n+6≥0xFF, then self_def[18:11]=0xFF, and an infinite number specified by IEEE 754 is output. The compression result is BF16_out

[15] =s, BF16_out[14:7]=8’hFF, and BF16_out[6:0]=7’h3F. If exp_max–n+6<0, a subnormal number specified by IEEE 754 is output. If there is a NaN specified by IEEE754, a quiet NAN is output, and the compression result is BF16_out

[15] =s, BF16_out[14:7]=8’hFF, and BF16_out[6:0]=7’h81.

[0141] The tensor core operation component provided by the present invention supports inputs in the FP16, BF16, FP8_e5m2 (1-5-2), and FP8_e4m3 (1-4-3) number systems and outputs in the FP16 and BF16 number systems. The supported number systems exactly match the number systems required by the large language model. The 16×16×16 tensor core multiply-add dimension in the FP16 number system and the 16×16×32 tensor core multiply-add dimension in the FP8 number system adopted by the present invention meet the matrix multiplication calculation dimension requirements of the large language model. Through customized design, it can cover the functional requirements without additional waste, thereby improving the chip performance and reducing the chip cost.

[0142] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A tensor core operation component, characterized in that It includes a decompression layer, a calculation layer, and a compression layer, where: The decompression layer is used to decompress the floating-point elements in the first matrix of a preset floating-point number system to obtain the first floating-point elements in a preset bit format, or to decompress the floating-point elements in the second matrix of the preset floating-point number system to obtain the second floating-point elements in the preset bit format, where the first matrix and the second matrix are two matrices in matrix multiplication operations; the preset floating-point number system includes a 16-bit floating-point number system and an 8-bit floating-point number system; The calculation layer is used to perform multiply-accumulate operations on all the first floating-point elements corresponding to the first matrix and all the second floating-point elements corresponding to the second matrix to obtain a floating-point number in a target bit format; The compression layer is used to receive the floating-point number in the target bit format and perform compression conversion processing on the floating-point number in the target bit format to obtain a floating-point operation result in a target floating-point number system; The decompression layer includes 512 identical decompression units; the preset bit format is a 20-bit format with bits from the 0th bit to the 19th bit, including a first preset bit format and a second preset bit format, where: When the preset floating-point number system is the 16-bit floating-point number system, the decompression unit is used to decompress the floating-point elements in one of the first matrices to obtain the first floating-point elements in the first preset bit format; or to decompress the floating-point elements in one of the second matrices to obtain the second floating-point elements in the first preset bit format; Among them, bits 19 to 12 in the first preset bit format are 8-bit exponents, and bits 11 to 0 in the first preset bit format are 12-bit signed mantissas; When the preset floating-point number system is the 8-bit floating-point number system, the decompression unit is used to decompress the floating-point elements in two of the first matrices to obtain the first floating-point elements in the second preset bit format; or to decompress the floating-point elements in two of the second matrices to obtain the second floating-point elements in the second preset bit format; Among them, bits 19 to 15 in the second preset bit format are 5-bit exponent bits, bits 14 to 10 in the second preset bit format are 5-bit signed mantissa bits, bits 9 to 5 in the second preset bit format are 5-bit exponent bits, and bits 4 to 0 in the second preset bit format are 5-bit signed mantissa bits.

2. The tensor core operation component according to claim 1, wherein The computing layer is composed of 256 identical computing units. The computing units include a first exponent addition module, a second exponent addition module, a first maximum exponent calculation module, a second maximum exponent calculation module, a first mantissa right shift bit number calculation module, a second mantissa right shift bit number calculation module, a first mixed radix mantissa multiplication module, a second mixed radix mantissa multiplication module, a mantissa alignment and accumulation module, and a normalization module, where: The first exponent addition module includes 16 groups of 8-bit adders, and the second exponent addition module includes 16 groups of 5-bit adders; The first mantissa right shift bit number calculation module includes 16 groups of 9-bit subtracters, and the second mantissa right shift bit number calculation module includes 16 groups of 6-bit subtracters; The first mixed radix mantissa multiplication module includes 16 groups of 12-bit by 12-bit signed number multipliers, and the second mixed radix mantissa multiplication module includes 16 groups of 5-bit by 5-bit signed number multipliers.

3. The tensor core operation component according to claim 2, wherein The first exponent addition module is used for: When the preset floating-point number system is the 16-bit floating-point number system, adding the 8-bit exponents in the first floating-point element in the first preset bit format and the 8-bit exponents in the second floating-point element in the first preset bit format, and subtracting the corresponding floating-point exponent offset value in the addition result to obtain a first exponent addition calculation result; When the preset floating-point number system is the 8-bit floating-point number system, adding the 5-bit exponents in the first floating-point element in the second preset bit format and the 5-bit exponents in the second floating-point element in the second preset bit format, and subtracting the corresponding floating-point exponent offset value in the addition result to obtain a second exponent addition calculation result; The second exponent addition module is used for: When the preset floating-point number system is the 16-bit floating-point number system, performing operand isolation and register clock gating; When the preset floating-point number system is the 8-bit floating-point number system, adding the 5-bit exponents in the first floating-point element in the second preset bit format and the 5-bit exponents in the second floating-point element in the second preset bit format, and subtracting the corresponding floating-point exponent offset value in the addition result to obtain a third exponent addition calculation result.

4. The tensor core operation component according to claim 3, wherein When the preset floating-point number system is the 16-bit floating-point number system, the first maximum exponent calculation module outputs a first maximum exponent, where the first maximum exponent is the maximum bit exponent in the first exponent addition calculation result output by the first exponent addition module; When the preset floating-point number system is the 8-bit floating-point number system, the first maximum exponent calculation module or the second maximum exponent calculation module outputs a second maximum exponent, where the second maximum exponent is the larger one between the first branch maximum exponent and the second branch maximum exponent, the first branch maximum exponent is the maximum bit exponent in the addition calculation result of the second exponents output by the first maximum exponent calculation module, and the second branch maximum exponent is the maximum bit exponent in the addition calculation result of the third exponents output by the second maximum exponent calculation module.

5. The tensor core operation component according to claim 4, wherein When the preset floating-point number system is the 16-bit floating-point number system, the first mantissa right shift bit number calculation module is used to output a first exponent difference, and the output result of the second mantissa right shift bit number calculation module is 0, where the first exponent difference is the exponent difference between the addition calculation result of the first exponents and the first maximum exponent; When the preset floating-point number system is the 8-bit floating-point number system, the first mantissa right shift bit number calculation module is used to output a second exponent difference, and the second mantissa right shift bit number calculation module is used to output a third exponent difference, where the second exponent difference is the exponent difference between the addition calculation result of the second exponents and the second maximum exponent; the third exponent difference is the exponent difference between the addition calculation result of the third exponents and the second maximum exponent.

6. The tensor core operation component according to claim 5, characterized in that, The first mixed number system mantissa multiplication module is used for: When the preset floating-point number system is the 16-bit floating-point number system, multiplying the 12-bit signed number mantissa in the first floating-point number element in the first preset bit format by the 12-bit signed number mantissa in the second floating-point number element in the first preset bit format to obtain a first mantissa multiplication calculation result; When the preset floating-point number system is the 8-bit floating-point number system, multiplying the 5-bit signed number mantissa in the first floating-point number element in the second preset bit format by the 5-bit signed number mantissa in the second floating-point number element in the second preset bit format to obtain a second mantissa multiplication calculation result; The second mixed number system mantissa multiplication module is used for: When the preset floating-point number system is the 16-bit floating-point number system, perform operand isolation and register clock gating; When the preset floating-point number system is the 8-bit floating-point number system, multiplying the 5-bit signed number mantissa in the first floating-point number element in the second preset bit format by the 5-bit signed number mantissa in the second floating-point number element in the second preset bit format to obtain a third mantissa multiplication calculation result.

7. The tensor core operation component according to claim 6, wherein The mantissa alignment and accumulation module is used for: When the preset floating-point number system is the 16-bit floating-point number system, according to the first exponent difference, the mantissa bits in the multiplication calculation results of multiple groups of the first mantissas are arithmetically right-shifted, and the multiplication calculation results of the first mantissas after the arithmetic right-shift of multiple groups of mantissa bits are aligned starting from the high bit, and the alignment results are accumulated to obtain the first mantissa alignment and accumulation result; When the preset floating-point number system is the 8-bit floating-point number system, according to the second exponent difference, the mantissa bits in the multiplication calculation results of multiple groups of the second mantissas are arithmetically right-shifted; according to the third exponent difference, the mantissa bits in the multiplication calculation results of multiple groups of the third mantissas are arithmetically right-shifted; and the multiplication calculation results of the second mantissas after the arithmetic right-shift of multiple groups of mantissa bits and the multiplication calculation results of the third mantissas after the arithmetic right-shift of multiple groups of mantissa bits are aligned starting from the high bit, and the alignment results are accumulated to obtain the second mantissa alignment and accumulation result.

8. The tensor core operation component according to claim 7, wherein The normalization module is used for: determining the target sign bit value according to the highest sign bit in the first mantissa alignment and accumulation result or the second mantissa alignment and accumulation result; converting the mantissa in the first mantissa alignment and accumulation result or the second mantissa alignment and accumulation result into an unsigned number, and obtaining the index value corresponding to the position of the first 1 in the unsigned number in the direction from the highest bit to the lowest bit; based on the index value, shifting the unsigned number to the left to obtain the target mantissa bits; constructing the target exponent bits according to the first maximum exponent and the index value; or, constructing the target exponent bits according to the second maximum exponent and the index value; obtaining the target bit format floating-point number according to the target sign bit value, the target exponent bits and the target mantissa bits; wherein, the target bit format floating-point number is a 20-bit floating-point number.

9. The tensor core operation component according to claim 8, wherein The compression layer includes 256 identical floating-point compression units, and the floating-point compression unit is used for: when the target floating-point number system is the FP16 number system, compressing the target bit format floating-point number into the floating-point operation result of the FP16 number system based on the IEEE754 standard; when the target floating-point number system is the BF16 number system, compressing the target bit format floating-point number into the floating-point operation result of the BF16 number system based on the IEEE754 standard; wherein, during the compression process, if the target exponent bits in the target bit format floating-point number are greater than or equal to the preset floating-point position, the floating-point compression unit outputs an infinite number; if the target exponent bits in the target bit format floating-point number are less than or equal to 0, the floating-point compression unit outputs a subnormal number; if the target bit format floating-point number is a NaN, the floating-point compression unit outputs a quiet NaN, where the preset floating-point position is the maximum value corresponding to the exponent bits of the target floating-point number system.

Citation Information

Patent Citations

  • Training neural network accelerators using mixed precision data formats

    CN113196305A