Split type coding tensor calculation framework and application thereof

By setting up a centralized encoder outside the tensor calculation unit and using a low-bit width encoding algorithm, the problems of encoder redundancy and data flow complexity in traditional tensor calculation units are solved, and higher area efficiency and energy efficiency are achieved, suitable for deep learning inference and high-performance computing.

CN120258048APending Publication Date: 2025-07-04HEFEI HENGSHUO SEMICON CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510392415.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

There are problems of encoder redundancy and complex data flow in traditional tensor computing units, resulting in increased chip area and power consumption, making it difficult to meet the real-time and energy efficiency requirements of deep learning and high-performance computing.

Method used

The split encoding tensor computing architecture is adopted, and the centralized encoder module is set outside the tensor computing unit microarchitectronics module, the multiplier is encoded through a low-bit width encoding algorithm, and the encoding results are broadcast or pulsated to the entire tensor computing unit microarchitectronics module, reducing repeated encoding operations and optimizing the data flow path.

Benefits of technology

It significantly reduces chip area and power consumption, improves computing efficiency, and is especially suitable for deep learning inference, edge computing and high-performance computing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258048A_ABST
    Figure CN120258048A_ABST
Patent Text Reader

Abstract

The invention relates to the field of neural network design, and discloses a split type coding tensor computing architecture and application, the architecture specifically comprises a tensor computing unit micro-architecture module and a centralized encoder module, the centralized encoder module is used for coding multiplicand, and the tensor computing unit micro-architecture module is used for coding the multiplicand. And the coding result is broadcasted or pulsated to the whole tensor calculation unit micro-architecture module, and the tensor calculation unit micro-architecture module is used for receiving the coding result and executing partial product compression and accumulation in multiplication operation to generate a final multiplication operation result. According to the method, the problems of encoder redundancy and complex data flow in a traditional AI accelerator are solved, the area efficiency and the energy efficiency are remarkably improved, and the method is particularly suitable for deep learning reasoning, edge calculation and high-performance calculation scenes and has practical value in practical significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural network design, and particularly to a split-type encoded tensor computing architecture and its application. Background Art

[0002] With the rapid development of artificial intelligence technology, its application scenarios have been widely penetrated into various fields of daily life and industrial production, such as search engines, autonomous driving, image recognition, natural language processing, etc. The core of these applications lies in processing and analyzing large-scale data, and tensor computing (especially matrix multiplication) is the basis for realizing these functions. For example, in deep learning, both the forward propagation and backward propagation processes of neural networks rely on a large number of matrix multiplication operations. With the continuous expansion of model scale (such as large language models LLMs), the demand for computing resources has increased sharply. Traditional general-purpose processors (such as CPUs) are inefficient in processing these large-scale tensor computations and are difficult to meet the requirements of real-time performance and energy efficiency. Therefore, developing dedicated hardware accelerators to improve the performance of tensor computing has become an inevitable choice.

[0003] On the one hand, in the existing TCU (Tensor Computing Unit) microarchitecture, each PE (Process Element) contains independent encoder logic, resulting in the same multiplicand being repeatedly encoded multiple times in large-scale matrix multiplication operations, causing waste of computing resources. From the perspective of hardware circuits, since each PE contains encoder logic, it leads to an increase in chip area and power consumption. Especially in a large-scale multiplier array, the repeated encoding logic not only increases the chip area but also increases the data transmission path length between PEs, further increasing power consumption; on the other hand, the existing encoding methods of multiplication algorithms will result in a relatively high encoding bit width, and a relatively high encoding bit width will increase the number of wire networks inside the PE, thereby having a negative impact on area, power consumption, etc. Therefore, both aspects of the problem need to be solved urgently. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a split-type encoded tensor computing architecture, which not only reduces the chip area and power consumption, but also optimizes the data flow path, significantly improving the computing efficiency.

[0005] The present invention adopts the following technical solutions to solve the technical problems:

[0006] The present invention provides a split-type encoded tensor computing architecture, including a tensor computing unit microarchitecture module and a centralized encoder module, wherein,

[0007] The centralized encoder module is disposed outside the tensor calculation unit microarchitecture module and is configured to encode the multiplicand and broadcast or pulse the encoding result to the entire tensor calculation unit microarchitecture module;

[0008] The tensor calculation unit microarchitecture module is composed of a number of core calculation units and is configured to receive the encoding result and perform partial product compression and accumulation in the multiplication operation to generate the final multiplication operation result.

[0009] Preferably, the tensor calculation unit microarchitecture module adopts one of the microarchitectures including 2D-Matrix, 1D / 2DArray, Systolic Array, and 3D-Cube.

[0010] Preferably, the core calculation unit is composed of a partial product compressor and a full adder and is configured to perform step by step:

[0011] Generate partial products according to the received encoding result;

[0012] Compress the generated partial products to generate two rows of sum and carry data;

[0013] Accumulate the compressed row and carry data to generate the multiplication result.

[0014] Preferably, when the tensor calculation unit microarchitecture module adopts the 2D-Matrix or 1D / 2DArray microarchitecture, the centralized encoder module is disposed outside the broadcast dimension of the 2D-Matrix or 1D / 2DArray microarchitecture;

[0015] When the tensor calculation unit microarchitecture module adopts the SystolicArray microarchitecture, the centralized encoder module is disposed outside the pulse dimension of the SystolicArray microarchitecture;

[0016] When the tensor calculation unit microarchitecture module adopts the 3D-Cube microarchitecture, the centralized encoder module is disposed outside one operand dimension of the 3D-Cube microarchitecture.

[0017] Preferably, the encoding algorithm adopted by the centralized encoder module includes the following steps:

[0018] Obtain the binary original code of the multiplicand A; the binary original code includes a sign bit sign and a numerical bit a i , where sign identifies the positive or negative of the multiplicand A, and a i is the i-th bit of the binary original code representation of the multiplicand A;

[0019] Generate intermediate encoding based on the binary original code using a low-bit-width polynomial;

[0020] The two-bit binary number is mapped to a low-bit coefficient compression code based on the intermediate code, and the number of bits of the compression code is the number of bits of the binary original code plus 1.

[0021] Preferably, the low-bit width polynomial satisfies:

[0022]

[0023] where |A| is the unsigned value of the multiplicand A, m is the number of bits after mating the binary coding bits m1 of the unsigned value of the multiplicand A, and w i is the value of the i-th bit of the intermediate code, and the w i is configured to generate four different values through a recursive expression and a carry chain code;

[0024] The m is the number of bits after mating the number of bits m1 of the multiplicand A value, specifically including:

[0025] If m1 is even, then m = m1, and the binary original code of the multiplicand A is the binary coding of the unsigned value of the multiplicand A;

[0026] If m1 is odd, then m = m1 + 1, and the binary original code of the multiplicand A is obtained by padding 0 in front of the highest bit of the binary coding of the unsigned value of the multiplicand A.

[0027] Preferably, the w i is configured to generate specifically including:

[0028] Generate a carry symbol C i using a carry chain code, and generate w i using a recursive expression, and the logical calculation expression is:

[0029] c i+1 =(a 2i+1 &a 2i )|(a 2i+1 &C i )

[0030] where c0 = 0, a 2i+1 , a 2i , a 2i+1 are all the values of the multiplicand A, and i ≥ 0;

[0031] w i =[a 2i+1 a 2i 10 +C i

[0032] [a 2i+1 a 2i 10Represents a 2-bit binary number a 2i+1 a 2i The decimal number of, w i ∈ {3, 0, 1, 2};

[0033] The middle encoding uses two-bit binary numbers to map to low-bit coefficient compression encoding, specifically:

[0034] The values {0, 1, 2, 3} of the middle encoding are represented by 2-bit binary {00, 01, 10, 11} respectively;

[0035] Then the 2-bit binary values {00, 01, 10, 11} after representation are mapped to the values {0, 1, 2, -1} of each bit K i Of the compression encoding to generate the compression encoding;

[0036] The bit K of the compression encoding i The bit weight of is 2 2i .

[0037] The present invention also provides a system-on-chip for performing neural network inference operations, including a three-level storage module, a controller module, a vector processing engine, and a tensor processing engine. The tensor processing engine is designed using the split encoding tensor calculation architecture as described above.

[0038] Preferably, the three-level storage module includes an external storage unit, a global cache unit, and an activation weight cache unit respectively, where

[0039] The external storage unit is configured to store large-scale neural network models and data;

[0040] The global cache unit is configured to cache the data read from the external storage unit for use by the tensor processing engine and the vector processing engine;

[0041] The activation weight cache unit is configured to store the activation values and weight data of the neural network to support efficient data reuse;

[0042] The controller module is configured to control read and write operations and preprocess convolution operations;

[0043] The tensor processing engine module includes a tensor calculation unit microarchitecture module and a centralized encoder module. The centralized encoder module is arranged in the read path of the weight cache and is used to encode the weight data into a format suitable for calculation by the tensor calculation unit microarchitecture module. The tensor calculation unit microarchitecture module is responsible for performing matrix multiplication and convolution operations;

[0044] The vector processing engine includes a number of arithmetic logic units and is configured to perform quantization, pooling, scalar addition, and activation function operations.

[0045] The present invention also provides an electronic device, comprising:

[0046] At least one data input port;

[0047] At least one data output port; and

[0048] The system-on-chip as described above.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] Through the low-width encoding algorithm and the externally encoded TCU microarchitecture, the present invention solves the problems of encoder redundancy and complex data flow in traditional AI accelerators, significantly improving the area efficiency and energy efficiency. The synergistic effect of these two technical points enables the encoding method and architecture in the present invention to show higher practical value in mainstream TCU microarchitectures (such as 2D-Matrix, SystolicArray, etc.), and is particularly suitable for deep learning inference, edge computing, and high-performance computing scenarios.

[0051] Regarding other outstanding substantive features and remarkable progress of the present invention compared with the prior art, they will be further described in detail in the embodiment part. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] By reading the following detailed description of the non-limiting embodiments with reference to the accompanying drawings, other features, objects, and advantages of the present invention will become more apparent:

[0053] Figure 1 Structural diagrams of five TCU computing microarchitectures in the prior art;

[0054] Figure 2 Schematic diagram of the algorithm flow under the prior art architecture;

[0055] Figure 3 Schematic diagram of the connection structure and signal processing of the encoder and selector in the prior art;

[0056] Figure 4 Schematic diagram of the MBE encoding algorithm logic in the prior art;

[0057] Figure 5 Structural diagrams of five split-type encoded tensor computing architectures in Embodiment 1;

[0058] Figure 6 Schematic diagram of the algorithm logic of the split-type encoded tensor computing architecture in Embodiment 1;

[0059] Figure 7 Schematic diagram of the hardware structure mapped by the encoding algorithm adopted by the split-type encoded tensor computing architecture in Embodiment 1;

[0060] Figure 8 Schematic diagram of the encoding algorithm when the multiplicand A is 91 in Embodiment 1;

[0061] Figure 9 Schematic diagram of the encoding algorithm when the multiplicand A is 124 in Embodiment 1;

[0062] Figure 10 Schematic diagram for comparing the encoding algorithm of the present invention with the prior art algorithm when the multiplicand A is 34 in Embodiment 1;

[0063] Figure 11 Schematic diagram for comparing the encoding algorithm of the present invention with the prior art algorithm when the multiplicand A is 50 in Embodiment 1;

[0064] Figure 12 Schematic diagram for comparing the power consumption of 5 structures in this embodiment with the prior art under different calculation digits;

[0065] Figure 13 Schematic diagram for comparing the area and power consumption of 5 structures in this embodiment with the prior art;

[0066] Figure 14 Schematic diagram for comparing the improvement of the area efficiency and energy efficiency of the present invention when the array scale is enlarged;

[0067] Figure 15 Schematic diagram of the structure of the system-on-chip in Embodiment 2. Detailed implementation manners

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0069] It should be noted that certain names are used to refer to specific components in the specification and claims. It should be understood that those of ordinary skill in the art may use different names to refer to the same component. The specification and claims of this application do not use the difference in names as a way to distinguish components, but use the substantial difference in functions of components as the criterion for distinguishing components. For example, the "comprising" or "including" used in the specification and claims of this application is an open-ended term, which should be interpreted as "including but not limited to" or "including but not limited to". The embodiments described in the detailed implementation manners section are the preferred embodiments of the present invention and are not used to limit the scope of the present invention.

[0070] Embodiment 1

[0071] Before introducing the technical solution of the present invention, in order to compare the prominent substantive features and remarkable progress of the present invention with the prior art, the prior art will be further explained and described.

[0072] On the one hand, as Figure 1 shown, a typical existing TCU computing microarchitecture includes:

[0073] 2D-Matrix microarchitecture ( Figure 1 (a)), the 2D-Matrix microarchitecture shown in Figure 1 (a) was proposed in the literature "Flexflow: A flexible dataflow accelerator architecture for convolutional neural networks". In each clock cycle, the operands are broadcast in each row and column, and the multiply-accumulate operation of the operands is calculated in each PE;

[0074] 1D / 2D Array microarchitecture ( Figure 1 (b)), the 1D / 2D Array microarchitecture shown in Figure 1 (b) was proposed in the literature "Trapezoid: A versatile accelerator for dense and sparse matrix multiplications". Based on the computing hardware of the addition tree, the operands in a row of the multiplicand matrix are broadcast to different computing arrays to parallelly calculate the vector inner product of a row and multiple columns of the matrix;

[0075] 2D-Systolic (OS) microarchitecture ( Figure 1 (c)), 2D-Systolic (WS) microarchitecture ( Figure 1 (d)), the 2D-Systolic (OS and WS) microarchitecture shown in Figure 1 (c) and (d) was proposed in the literature "Tpuv4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings". Among them, OS (Output Stationary) adopts the output fixed data flow of the calculation result, and its operands pulsate in each row and column of the computing array; WS (Weight Stationary) adopts the data flow with the weights fixed locally in the PE, and its multiplicand pulsates in each column of the PE, while the multiplier (weights) are stored inside each PE;

[0076] 3D-Cube microarchitecture (Figure 1 (e)), in the literature "Ascend: a scalable and unified architecture for ubiquitous deep neural network computing: Industry track paper", the 3D-Cube microarchitecture as shown in Figure 1 (e) is proposed, where two matrix operands and partial sum results are pipelined in different dimensions of the PE array respectively.

[0077] These computing architectures use the multiply-accumulator (MAC) as the basic processing element (PE), and optimize the data flow path through different spatial interconnection topologies to achieve efficient matrix multiplication computing functions.

[0078] As Figure 2 shown, when considering a single multiplication operation, first, the encoding part is only related to the logical calculation related to the multiplicand A and has nothing to do with the multiplier B; second, the result of the multiplication is directly related to the encoded multiplicand A and the multiplier B.

[0079] When this behavior is extended to the Figure 1 computing array in, due to the phenomenon of multiplying the same multiplicand A by different multipliers B in the spatial or temporal dimension in matrix multiplication or convolution, and in the Figure 1 computing architecture in, it is manifested as data broadcasting and pipelining of operands, so there will be repeated encoding behavior of the multiplicand A inside the PE (as shown in Figure 2 (a)). From the perspective of the TCU, what the PE needs is the encoded multiplicand A, not the original A. When applied to various existing TCU hardware architectures, whether it is the 2D-Matrix or 1D / 2D Array based on data broadcasting, or the Systolic Array or 3D Cube based on data flow, as Figure 2 (b) shown, there is repeated encoding behavior for the same operands in traditional tensor computing units. Therefore, this process has redundancy of encoders in hardware for the TCU, increasing the logical area and computing power consumption of the TCU, and at the same time increasing the logical delay inside the PE.

[0080] This shows that in the existing TCU microarchitecture, each PE contains independent encoder logic, resulting in the same multiplicand being repeatedly encoded multiple times in large-scale matrix multiplication operations, causing waste of computing resources. From the perspective of the hardware circuit, since each PE contains encoder logic, it leads to an increase in chip area and power consumption. Especially in a large-scale multiplier array, the repeated encoding logic not only increases the chip area, but also increases the data transmission path length between PEs, further increasing the power consumption.

[0081] On the other hand, most encoders in modern multipliers use the MBE encoding method. For an n-bit operand A(a n-1 a n-2 …a0), where a i is the i-th bit of the two's complement representation of A. The MBE encoder represents A in the following form:

[0082]

[0083] where a -1 = 0, m i ∈{-2, -1, 0, 1, 2}. Therefore, when A×B is executed inside the multiplier, as Figure 3 shown, the selector after the encoder will select -2B, -B, 0, B, and 2B to the subsequent adder unit according to the signals NEG, CE, and SE encoded from A. The logical expressions of NEG, CE, and SE are as follows:

[0084]

[0085] For example, when A = 34, as Figure 4 shown, after the MBE encoding logic, four encoding coefficients with different bit weights {2 6 ,2 4 ,2 2 ,2 0} are obtained. Multiplying the encoding coefficients with different bit weights by B and then summing them can obtain the multiplication result of A×B. This process only requires a hardware circuit of shift and addition to achieve multiplication calculation. For an n-bit two's complement number, the MBE encoding method will obtain the encoding value of the bits( Figure 4 the MBE encoding bit width of INT8 in

[0086] is 12 bits). Therefore, the higher encoding bit width will increase the number of wire networks inside the PE, which will have a negative impact on the area, power consumption, etc. Specifically, the essence of MBE encoding is to encode every 2 bits of the multiplier into 3-bit control signals (NEG, SE, CE). For the multiplication of two n-bit numbers, the multiplicand needs to be encoded into

[0087] bits. This encoding method significantly increases the data width after encoding. The increase in the data width after encoding will cause the interconnection lines for transmitting these data to become wider. In chip design, the increase in the width of the interconnection lines will occupy more chip area and increase the power consumption of signal transmission. This is not conducive to the area optimization and power control of the chip.

[0088] This embodiment provides a split-type encoded tensor calculation architecture, including a tensor calculation unit microarchitecture module and a centralized encoder module. Among them,

[0089] The centralized encoder module is arranged outside the tensor calculation unit microarchitecture module, configured to encode the multiplicand, and broadcast or pulse the encoding result to the entire tensor calculation unit microarchitecture module;

[0090] In this embodiment, the tensor calculation unit is an architecture module that adopts one of 2D-Matrix, 1D / 2DArray, Systolic Array, and 3D-Cube microarchitectures. Generally, 2D-Matrix and 1D / 2DArray are suitable for processing medium-scale matrix multiplication and convolution operations, and are widely used in AI accelerators in edge computing and mobile devices. SystolicArray (OS and WS) is suitable for processing large-scale deep learning tasks, such as NLP and computer vision, by optimizing the data flow path. The 3D-Cube architecture further improves the computing parallelism by adding a third dimension, and is suitable for high-performance computing and the training and inference of ultra-large-scale deep learning models. Those skilled in the art can choose according to their needs and will not be elaborated here.

[0091] The tensor calculation unit microarchitecture module is composed of several core calculation units, configured to receive the encoding result, and perform partial product compression and accumulation in the multiplication operation to generate the final multiplication operation result.

[0092] In this embodiment, the core calculation unit is composed of a partial product compressor and a full adder, configured to perform step by step:

[0093] Generate partial products according to the received encoding result;

[0094] Compress the generated partial products to generate two rows of sum and carry data;

[0095] Accumulate the compressed row and carry data to generate the multiplication result.

[0096] In this embodiment, when the tensor calculation unit microarchitecture module adopts a 2D-Matrix or 1D / 2DArray microarchitecture, the centralized encoder module is arranged outside the broadcast dimension of the 2D-Matrix or 1D / 2DArray microarchitecture;

[0097] When the tensor calculation unit microarchitecture module adopts a Systolic Array microarchitecture, the centralized encoder module is arranged outside the pulse dimension of the SystolicArray microarchitecture;

[0098] When the tensor computing unit microarchitecture module adopts the 3D-Cube microarchitecture, the centralized encoder module is set outside one operand dimension in the 3D-Cube microarchitecture.

[0099] Specifically, as Figure 5 shown, first, the encoder logic inside all PEs in these architectures is removed, which is the core computing unit (RPE). Among them, an encoder is placed outside the broadcast dimension of the 2D-Matrix and 1D / 2DArray architectures to encode the operand and then broadcast it; for the SystolicArray, an encoder is placed outside the systolic dimension of the multiplicand to encode the operand and then input it into the computing array for encoded number pulsation; while the 3D-Cube architecture is a three-dimensional tensor computing architecture, usually composed of multiple layers of PEs stacked to form a cube structure. Data flows in three dimensions and can process data in parallel at multiple levels. This architecture further improves the computing parallelism and data reuse ability by adding a third dimension. Therefore, an encoder matrix is placed outside one operand dimension to encode the operand, enabling the encoded numbers to flow in the cube, reducing encoder redundancy, optimizing the data flow path, reducing chip area and power consumption, and improving computing efficiency.

[0100] As Figure 6 shown, the core module of the present invention consists of two parts: the tensor computing unit microarchitecture module and the centralized encoder module. These two parts work together to achieve efficient matrix multiplication operations by reducing encoder redundancy and optimizing the data flow path. The computing array in the figure can be replaced with other TCU microarchitectures, which will not be elaborated here. The following is a detailed description of these two parts:

[0101] The RPE (Reduced Processing Element) is the core computing unit in the present invention, which is characterized by removing the encoder logic in the traditional PE. In the traditional multiplication unit, each PE contains an independent encoder for encoding the multiplicand (A) to generate partial products. However, this design leads to a large number of repeated encoding operations. Especially in matrix multiplication, the same multiplicand will be encoded multiple times, resulting in a waste of computing resources. In the RPE, the encoder logic is removed, and only the partial product compressor (Compressor Tree) and the full adder (Full Adder) are retained. This design reduces the area and power consumption of each PE, while simplifying the data flow path. The main task of the RPE is to perform partial product compression and accumulation in the multiplication operation without undertaking the encoding task. This division of labor enables the RPE to complete the computing task more efficiently. The specific advantages can be reflected in the reduction of chip area. Due to the removal of the encoder logic, this area optimization effect is more obvious in large-scale multiplier arrays. At the same time, removing the encoder logic not only reduces the static power consumption but also shortens the data transfer path between PEs, further reducing the dynamic power consumption.

[0102] To ensure the basic function of multiplication, the present invention introduces a centralized encoder module outside the RPE array. This encoder module is responsible for encoding the multiplicand (A) and broadcasting or pulsing the encoded result to the entire RPE array. Each RPE only needs to receive the encoded multiplicand without performing its own encoding operation. Through centralized low-width encoding, the entire RPE array only requires a small number of encoders instead of one encoder for each PE. This design significantly reduces the chip area and power consumption.

[0103] The RPE and the centralized encoder module cooperate to achieve efficient tensor calculation. The specific workflow is as follows:

[0104] Step 1: Encoding: The external encoder encodes the multiplicand (A) to generate the encoded result (Encoded(A));

[0105] Step 2: Partial product generation: The encoded multiplicand (Encoded(A)) is broadcast to the entire RPE array, and each RPE generates partial products according to the received encoded result;

[0106] Step 3: Partial product compression: The partial product compressor (Compressor Tree) in the RPE compresses the generated partial products to generate two rows of results (Sum and Carry);

[0107] Step 4: Accumulation: The full adder in the RPE accumulates the compressed results to generate the final multiplication and accumulation result.

[0108] The encoding algorithm adopted by the centralized encoder module in this embodiment includes the following steps:

[0109] Obtain the binary original code of the multiplicand A, which includes a sign bit sign and a numerical bit a i , where sign identifies the positive or negative of the multiplicand A, and a i is the i-th bit of the binary original code representation of the multiplicand A.

[0110] As Figure 7 shown, generate intermediate codes based on the binary original code using a low-bit-width polynomial; the low-bit-width polynomial in this embodiment satisfies:

[0111]

[0112] where |A| is the unsigned value of the multiplicand A, m is the number of bits after mating the binary coding bits m1 of the unsigned value of the multiplicand A, and w i is the value of the i-th bit of the intermediate code, and the w i is configured to generate four different values through a recursive expression and a carry chain coding.

[0113] In this embodiment, m is the number of bits after mating the numerical bits m1 of the multiplicand A, specifically including:

[0114] If m1 is even, then m = m1, and the binary original code of the multiplicand A is the binary coding of the unsigned value of the multiplicand A;

[0115] If m1 is odd, then m = m1 + 1, and the binary coding of the unsigned value of the multiplicand A is the binary original code of the multiplicand A after adding 0 in front of the highest bit.

[0116] In this embodiment, the w i is configured to generate specifically through a recursive expression and a carry chain coding, including:

[0117] Generate a carry symbol C i using a carry chain coding, and generate w i using a recursive expression. The logical calculation expression is:

[0118] c i+1 =(a 2i+1 &a 2i )|(a 2i+1 &C i )

[0119] where c0 = 0, a 2I+1, a 2i , a 2i+1 Both are the values of the multiplicand A, where i ≥ 0;

[0120] w i = [a 2i+1 a 2i 10 + C i

[0121] [a 2i+1 a 2i 10 represents the decimal number of the 2-bit binary a 2i+1 a 2i and w i ∈ {3, 0, 1, 2}.

[0122] Based on the intermediate coding, a two-bit binary number is mapped to a low-bit coefficient compression coding, and the number of bits of the compression coding is the number of bits of the binary original code plus 1. Specifically:

[0123] The values {0, 1, 2, 3} of the intermediate coding are represented by the 2-bit binary {00, 01, 10, 11} respectively;

[0124] Then, the 2-bit binary values {00, 01, 10, 11} after representation are mapped to the values {0, 1, 2, -1} of each bit K i of the compression coding to generate the compression coding;

[0125] In this embodiment, the bit weight of the compression coding bit K i is 2 2i .

[0126] The multiplier coding algorithm in this embodiment further includes:

[0127] Identifying the sign bit sign value of the multiplicand A:

[0128] If the sign value is negative, the sign bit of the multiplier B is inverted, otherwise the sign bit of the multiplier B remains unchanged;

[0129] Associating the bit weight 2 i with the value of the compression coding bit K 2i and then multiplying each by the multiplier to obtain the partial product;

[0130] Adding up the partial products to obtain the final multiplication operation result;

[0131] Among them, the operation of the partial product is implemented by a shift circuit, and the accumulation of the partial products is implemented by a register circuit and a full adder circuit.

[0132] To further clarify and verify the principle and performance of the multiplier coding algorithm of the present invention, the following is an example:​​

[0133] As Figure 8 shown:

[0134] When A = 91(0,1011011), the sign flag is 0, indicating that the value is positive;

[0135] After padding 0s at the high positions, its binary original code is (01011011);

[0136] Based on the binary original code (01011011), according to the calculation expression of w i the intermediate encoding value of A is represented as {0,1,2,3,3};

[0137] Using 2-bit binary {00,01,10,11} to represent the intermediate encoding value {0,1,2,3,3} respectively, we get {00,01,10,11,11};

[0138] Based on the represented {00,01,10,11,11}, it is mapped to the compressed encoding {1,2,-1,-1};

[0139] The corresponding bit weights are {2 6 ,2 4 ,2 2 ,2 0};

[0140] Multiplying the values of the compressed encoding bits K i associated with the bit weight 2 2i by the multiplier respectively to obtain the partial products, that is, A×B = 64B + 32B - 4B - B = 91B.

[0141] Again, as Figure 9 shown:

[0142] When A = 124(0,1111100), according to the calculation expression of w i the encoding value of A is represented as {0,2,0,-1,0}, so A×B = 128B - 4B = 124B.

[0143] To further compare the differences between the present invention and the prior art, the following comparison is continued

[0144] As Figure 10 shown:

[0145] When A = 34(00100010), after the MBE encoding logic, as Figure 10 (left) shown, the encoding value of A is {1, -2, 1, -2}, since the coefficient m of MBE i∈{-2, -1, 0, 1, 2}, so 3 bits are needed to represent each encoded bit. Therefore, a total of 12 encoded bits are required to represent the encoded value of A for the INT8 operand. Finally, the calculation formula of MBE is A×B = 64B - 32B + 4B - 2B = 34B. The encoding method proposed in this embodiment is as Figure 10 (right), according to the calculation expression of w i , the encoded value of A is {0, 0, 2, 0, 2}, and the calculation formula A×B = 32B + 2B = 34B. Only 9 encoded bits are needed to represent the encoded value of A.

[0146] When A = 50(00110010), after passing through the MBE encoding logic, as Figure 11 (left) shows, the encoded value of A is {1, -1, 1, -2}, and the final calculation formula of MBE is A×B = 64B - 16B + 4B - 2B = 50B. The encoding method proposed in this embodiment, as Figure 11 (right), according to the calculation expression of w i , the encoded value of A is {0, 1, -1, 0, 2}, and the calculation formula A×B = 64B - 16B + 2B = 50B. Since w i ∈{-1, 0, 1, 2}, compared with MBE, only 2 bits are used to represent one of the encoded bits.

[0147] Therefore, it can be undoubtedly concluded that for the multiplication of two n-digit numbers, the multiplicand in this embodiment needs to be encoded into n + 1 bits. Compared with the n-bit MBE method, the new design can significantly reduce the data width after encoding, thereby reducing the width and power consumption of the interconnection lines.

[0148] To further verify the superiority of the present invention, the circuit structure of the present invention was simulated and verified, as Figure 12 shown:

[0149] 2D-Matrix architecture: The energy consumption is reduced by 15.1% - 15.9%. Systolic Array (OS) architecture: The energy consumption is reduced by 11.3% - 12.8%. Systolic Array (WS) architecture: The energy consumption is reduced by 10.2% - 11.7%. 1D / 2D Array architecture: The energy consumption is reduced by 14.0% - 16.0%. 3D-Cube architecture: The energy consumption is reduced by 5.0% - 6.0%. It can be verified that through the centralized encoder design, the transmission path length of data between PEs is reduced, and the dynamic power consumption is further reduced.

[0150] The area and power consumption performance are as Figure 13As shown, in the 1D / 2D Array architecture, the layout optimization brought about by the reduction in PE area has increased the area efficiency by up to 20.2% (in the 1 TOPS scenario). In the Systolic Array and 3D Cube architectures, the area overhead is further reduced through data line width compression. For power consumption performance, in the 2D-Matrix architecture, the energy consumption is reduced by 15.1% - 15.9%; in the 1D / 2D array, it is reduced by 14.0% - 16.0%. The external encoder reduces the data transmission distance between PEs, thereby reducing the dynamic power consumption.

[0151] The architecture in this embodiment can be seamlessly integrated into existing TCU microarchitectures (such as 2D-Matrix, SystolicArray, 3D-Cube), and as the array scale expands (such as from 256 GOPS to 4 TOPS, as Figure 14 shown), the improvement in area efficiency and energy efficiency is further amplified (the area efficiency increases from 8.7% to 12.2%, and the energy efficiency increases from 13.0% to 17.5%).

[0152] In summary, in terms of technical effects, the area efficiency improvement of the TCU in this embodiment significantly reduces the area of a single PE by removing the encoder logic in the PE. In a large-scale multiplier array, this area optimization effect is even more obvious. In addition, the number of encoders is reduced: in the TCU computing array, the centralized encoder design proposed in this embodiment significantly reduces the number of encoders. For example, a 32×32 two-dimensional array only requires 32 encoders, saving 992 encoders.

[0153] The present invention realizes efficient matrix multiplication operations by removing the encoder logic from the RPE and centrally placing it outside the multiplier array. The RPE focuses on partial product compression and accumulation, while the external encoder is responsible for encoding and broadcasting the multiplicand. At the same time, the low-bit-width encoding algorithm ensures low overhead for encoding value broadcasting. This design not only reduces the chip area and power consumption but also optimizes the data flow path, significantly improving the computing efficiency. The present invention is applicable to various application scenarios such as deep learning chips, high-performance computing, and edge computing devices.

[0154] Embodiment 2

[0155] As Figure 15 shown, this embodiment provides a system-on-chip for performing neural network inference operations (such as ResNet34, ResNet50, ResNet101, InceptionV3, DenseNet121, DenseNet161, Vgg13, Vgg19, etc.), including a three-level storage module, a controller module, a vector processing engine, and a tensor processing engine. The tensor processing engine is designed using the split encoding tensor computing architecture as in Embodiment 1.

[0156] The three - level storage module in this embodiment respectively includes an external storage unit (DRAM), a global buffer unit (Global Buffer), and an activation weight buffer unit. Among them,

[0157] The external storage unit is configured to store large - scale neural network models and data;

[0158] The global buffer unit is configured to cache the data read from the external storage unit for use by the tensor processing engine and the vector processing engine;

[0159] The activation weight buffer unit (ActivationBuffer and WeightBuffer) is configured to store the activation values and weight data of the neural network, supporting efficient data reuse;

[0160] The controller module (Controller) is responsible for controlling the read - write operations of the SRAM and includes an img2col module for pre - processing convolution operations

[0161] The tensor processing engine (TensorProcessingEngine) includes a tensor calculation unit micro - architecture module and a centralized encoder module. The centralized encoder module is arranged in the read path of the weight buffer and is used to encode the weight data into a format suitable for the calculation of the tensor calculation unit micro - architecture module. The tensor calculation unit micro - architecture module is responsible for performing matrix multiplication and convolution operations;

[0162] The vector processing engine (SIMD VectorProcessing Engine) contains a number of arithmetic logic units and is configured to perform quantization, pooling, scalar addition, and activation function operations.

[0163] Since the SoC contains a large amount of on - chip SRAM, controllers, and SIMD vector processing engines, the area ratio of the encoder in the computing module is relatively low. Therefore, from the overall perspective of the SoC, the area benefit brought by the architecture proposed in this embodiment is relatively small, and the main advantage is to reduce the inference power consumption.

[0164] Embodiment 3

[0165] This embodiment provides an electronic device, including:

[0166] At least one data input port;

[0167] At least one data output port; and

[0168] The system - on - chip as described in Embodiment 2.

[0169] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0170] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A split-type encoded tensor calculation architecture, characterized in that, It includes a tensor calculation unit microarchitecture module and a centralized encoder module. Among them, the centralized encoder module is arranged outside the tensor calculation unit microarchitecture module, and is configured to encode the multiplicand and broadcast or pulse the encoding result to the entire tensor calculation unit microarchitecture module; the tensor calculation unit microarchitecture module is composed of a number of core calculation units, and is configured to receive the encoding result and perform partial product compression and accumulation in the multiplication operation to generate the final multiplication operation result.

2. The split coding tensor calculation architecture according to claim 1, wherein The tensor calculation unit microarchitecture module adopts one of the microarchitectures including 2D-Matrix, 1D / 2DArray, SystolicArray, and 3D-Cube.

3. The split coding tensor calculation architecture according to claim 2, characterized in that The core calculation unit is composed of a partial product compressor and a full adder, and is configured to perform step by step: generate partial products according to the received encoding result; compress the generated partial products to generate two rows of sum and carry data; accumulate the compressed row and carry data to generate the multiplication result.

4. The split coding tensor calculation architecture according to claim 3, wherein When the tensor calculation unit microarchitecture module adopts the 2D-Matrix or 1D / 2D Array microarchitecture, the centralized encoder module is arranged outside the broadcast dimension of the 2D-Matrix or 1D / 2DArray microarchitecture; When the tensor calculation unit microarchitecture module adopts the SystolicArray microarchitecture, the centralized encoder module is arranged outside the pulse dimension of the SystolicArray microarchitecture; When the tensor calculation unit microarchitecture module adopts the 3D-Cube microarchitecture, the centralized encoder module is arranged outside one operand dimension of the 3D-Cube microarchitecture.

5. The split-type encoded tensor calculation architecture according to claim 4, wherein The encoding algorithm adopted by the centralized encoder module includes the following steps: Obtain the binary original code of the multiplicand A; the binary original code includes a sign bit sign and a numerical bit a i , where sign identifies the positive or negative of the multiplicand A, and a i is the i-th bit of the binary original code representation of the multiplicand A; generate intermediate encoding based on the binary original code using a low-bit-width polynomial; map the intermediate encoding to a low-bit coefficient compression encoding using two-bit binary numbers, and the number of bits of the compression encoding is the number of bits of the binary original code plus 1.

6. The split coding tensor calculation architecture according to claim 5, characterized in that The low-bit-width polynomial satisfies: where |A| is the unsigned value of the multiplicand A, m is the number of bits after mating the binary coding bits m1 of the unsigned value of the multiplicand A, and w i is the value of the i-th bit of the intermediate code, and the w i is configured with four different values generated by a recursive expression and a carry chain code; The m is the number of bits after the number of bits m1 of the multiplicand A is paired. Specifically, it includes: If m1 is an even number, then m = m1, and the binary original code of the multiplicand A is the binary encoding of the unsigned value of the multiplicand A; If m1 is an odd number, then m = m1 + 1, and the binary encoding of the unsigned value of the multiplicand A with a 0 added in front of the highest bit is the binary original code of the multiplicand A.

7. A split-type encoded tensor calculation architecture according to claim 6, characterized in that The said w i configured to generate through recursive expressions and carry chain encoding specifically including: Generate carry symbol C using carry chain encoding i , generate w using recursive expression i , the logical calculation expression is: c i+1 = (a 2i+1 & a 2i ) | (a 2i+1 & C i ) where c0 = 0, a 2i+1 , a 2i , a 2i+1 are all the values of the multiplicand A, and i ≥ 0; w i = [a 2i+1 a 2i 10 + C i ​ [a 2i+1 a 2i 10 Represents the decimal number of the 2-bit binary a 2i+1 a 2i where w i ∈ {3, 0, 1, 2};​ The mapping of the intermediate encoding to a low-bit coefficient compression encoding using two-bit binary numbers is specifically: represent the values {0, 1, 2, 3} of the intermediate encoding using 2-bit binary {00, 01, 10, 11} respectively; Then, map the represented 2-bit binary values {00, 01, 10, 11} to the values {0, 1, 2, -1} of each bit K of the compression code to generate the compression code; i ​ The compressed coding bit K i has a bit weight of 2 2i .

8. A system-on-chip for performing neural network inference operations, characterized in that, It includes a three-level storage module, a controller module, a vector processing engine, and a tensor processing engine. The tensor processing engine adopts the split encoding tensor calculation architecture design described in any one of claims 1-7.

9. A system-on-chip according to claim 8, characterized in that, The three-level storage module respectively includes an external storage unit, a global cache unit, and an activation weight cache unit. Among them, the external storage unit is configured to store large-scale neural network models and data; A global cache unit configured to cache data read from the external storage unit for use by the tensor processing engine and the vector processing engine; An activation weight cache unit configured to store activation values and weight data of a neural network to support efficient data reuse; The controller module configured to control read and write operations and preprocess convolution operations; The tensor processing engine includes a tensor computing unit microarchitecture module and a centralized encoder module. The centralized encoder module is disposed in the read path of the weight cache and is configured to encode weight data into a format suitable for calculation by the tensor computing unit microarchitecture module. The tensor computing unit microarchitecture module is responsible for performing matrix multiplication and convolution operations; The vector processing engine includes a number of arithmetic logic units configured to perform quantization, pooling, scalar addition, and activation function operations.

10. An electronic device, characterized in that, Comprising: At least one data input port; At least one data output port; And A system-on-chip as claimed in any one of claims 8-9.