Reconfigurable multiply-accumulate unit for neural network processor and processor

By designing a reconfigurable multiplication-accumulation unit, the problem in existing technologies that hardware is difficult to be compatible with diverse data types is solved, and efficient support for multiple data precisions and calculation types is achieved in neural network processors, thereby improving computing efficiency and flexibility.

CN120687065APending Publication Date: 2025-09-23INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510742119.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing reconfigurable computing unit designs are difficult to fully cover the diverse data types and computing types in neural network training and inference, and ignore the research on the reconfigurability of data paths, resulting in hardware deficiencies in performance and efficiency.

Method used

A reconfigurable multiplication-accumulation unit for neural network processors is designed, which includes multiple modules such as input preprocessing, XOR array, addition array, reconfigurable multiplier, normalization array and output post-processing module. It can flexibly support various data precisions and calculation types, including 4-bit, 8-bit, 16-bit, 32-bit and 64-bit data types, and supports multiplication, multiply-add and multiply-accumulate operations.

Benefits of technology

It efficiently supports multiple data precisions and calculation types under one architecture, improves the flexibility and computational efficiency of neural network training and reasoning, and can cover a variety of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120687065A_ABST
    Figure CN120687065A_ABST
Patent Text Reader

Abstract

According to the reconfigurable multiply-accumulate unit for the neural network processor and the processor, an efficient reconfigurable multiplier is constructed by organizing a large number of low-precision multiplication units. A peripheral reconfigurable computing circuit is designed to form a multiply-accumulate unit, so that the multiply-accumulate unit can flexibly support computing requirements of various precisions, various computing types and various data paths, and the compatibility and efficiency of hardware to a neural network are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of processor design for neural network training and reasoning in the field of machine learning, and in particular to a reconfigurable multiplication-accumulation unit for a neural network processor and a neural network processor. Background Art

[0002] With the rapid development of deep learning, the scale and complexity of neural networks continue to climb. From early small networks to today's large-scale deep models, networks have become deeper and more complex, with exponential growth in parameter size and computational complexity. This progress has significantly improved the performance of neural networks in solving complex tasks such as natural language processing and computer vision. However, it has also brought significant computational and storage pressures, resulting in significantly increased model training and inference times, and placing higher performance and energy efficiency requirements on hardware devices.

[0003] To address these issues, quantization technology has been widely adopted. By converting high-precision data in neural networks (such as 32-bit floating-point numbers) to lower-precision data (such as 16-bit, 8-bit, or even 4-bit integers), quantization significantly reduces computational and storage requirements, accelerating training and inference. However, the quantization accuracy residuals vary across neural network models and network layers, requiring hardware capable of multi-precision computing.

[0004] To this end, many studies have adopted reconfigurable concepts in the hardware circuit design of computing units. These designs employ hybrid designs with varying computational precisions, including the integration of multiple fixed-point systems, multiple floating-point systems, and multiple floating-point and fixed-point systems. These designs have alleviated the aforementioned issues to a certain extent, resulting in gains in system performance and area.

[0005] First, existing reconfigurable computing unit designs generally have limited support for precision types, making it difficult to fully cover the diverse data types potentially used in neural network training and inference. Second, existing research has primarily focused on data type reconfiguration, while neglecting the study of reconfigurable computing types and data paths. Summary of the Invention

[0006] In neural network processors, how to efficiently be compatible with and support the various data precision types involved in neural network model training and inference on a single hardware architecture, and meet the hardware design challenges brought about by these diverse computing requirements.

[0007] In view of the shortcomings of the prior art, the present invention proposes a reconfigurable multiplication and accumulation unit for a neural network processor, which includes: a multiplier, an accumulator;

[0008] The multiplier includes: a first input pre-processing module, an XOR array, an addition array, a reconfigurable multiplier, a normalization array, an exponent adjustment array and an output post-processing module;

[0009] The first input preprocessing module is configured to slice the multiplier and the multiplicand according to their data precision to obtain a plurality of data pairs; input the sign bit of each data pair into the XOR array, input the exponent bit of each data pair into the addition array, and input the mantissa bit or the valid bit of each data pair into the reconfigurable multiplier according to the data type of the data pair;

[0010] The XOR array includes a plurality of XOR units for receiving the sign bits of each data pair, performing XOR processing on the sign bits, and obtaining the sign bits of the multiplied data pairs;

[0011] The addition array includes a plurality of adders for receiving the exponent bits of each data pair, performing addition processing on the exponent bits, obtaining the exponent addition result of each data pair, and passing the result to the exponent adjustment array;

[0012] The reconfigurable multiplier includes a plurality of multipliers and a group of adders for performing fixed-point multiplication and floating-point mantissa multiplication operations to obtain a multiplication result;

[0013] The normalization array includes a plurality of shifters for receiving the multiplication result and passing the exponent adjustment value to the exponent adjustment module, and passing the normalized mantissa bits to the output post-processing module;

[0014] The exponent adjustment array includes an adder or a data selector for receiving the exponent adjustment value to adjust the exponential addition result and pass the adjustment result to the output post-processing module;

[0015] The output post-processing module is used to obtain a multiplication result according to the sign bit, the adjustment result and the mantissa bit;

[0016] The accumulator includes: a second input preprocessing module, an exponent comparison array, a shift array, an addition array, a normalization array, an exponent adjustment array, and an accumulation register;

[0017] The second input preprocessing module slices the logarithms of the data to be added, splits the slicing results according to data types, inputs the exponent bits obtained by the splitting into the exponent comparison array, inputs the sign bit and the mantissa bits into the shift array, and splices the significant bits or the sign bit and the significant bits into the shift array;

[0018] The exponent comparison array comprises a plurality of comparators for comparing the exponent bits and passing the maximum exponent bit to the shift array and the exponent adjustment array;

[0019] The shift array includes a plurality of shifters, which are used to receive data provided by the second input pre-processing module and shift the data according to the maximum exponent bit to obtain a processing result and pass it to the addition array;

[0020] The addition array includes a plurality of adders, which are used to perform addition operations on the mantissas of the fixed-point data and the signed floating-point data according to the processing results to obtain operation results.

[0021] The normalization array includes a plurality of shifters for receiving the operation result, passing the exponent adjustment value to the exponent adjustment array, and passing the normalized mantissa bits in the operation result to the accumulation register;

[0022] The exponent adjustment array comprises an adder for adjusting the maximum exponent bit according to the exponent adjustment value and passing the adjustment result to the accumulation register;

[0023] The accumulator register is used to cache and merge data to obtain the multiplication and addition result.

[0024] The reconfigurable multiplication-accumulation unit for a neural network processor further includes: an output selector; the output of the output selector is the output of the accumulator and the output of the multiplier, and is used to select the output of the accumulator or the output of the multiplier according to a control instruction;

[0025] The accumulator also includes a data selector; the input of the data selector is the multiplication and addition result output by the accumulator register and the addend input to the accumulator, and is used to select, according to the control instruction, whether the data input to the second input preprocessing module is the multiplication and addition result output by the accumulator register or the addend.

[0026] The reconfigurable multiply-accumulate unit for a neural network processor, wherein the first input preprocessing module is used to:

[0027] If the multiplier and the multiplicand have the same precision of x bits, the multiplicand and the multiplier are sliced ​​into segments of x bits each, and the corresponding segments after slicing constitute a pair of data, with a total of 64 / x pairs; if the input data pairs are floating-point type, the sign bit of each data pair is input into the XOR array, the exponent bit of each data pair is input into the addition array, and the mantissa bit of each data pair is input into the reconfigurable multiplier; if the input data pairs are unsigned fixed-point type, the valid bits of the data are directly input into the reconfigurable multiplier; if the input data pairs are signed fixed-point type, the sign bit is input into the XOR array, and the data complement is taken from the original code and then input into the reconfigurable multiplier.

[0028] The reconfigurable multiplication-accumulation unit for a neural network processor, wherein the output post-processing module is responsible for data merging and conversion, if the multiplier and the multiplicand are floating-point types, the sign bit provided by the XOR array, the exponent bit provided by the exponent adjustment array, and the mantissa bit provided by the normalization array are spliced ​​to generate the results of each floating-point multiplication; if the multiplier and the multiplicand are signed fixed-point types, the valid bits provided by the normalization array are complemented according to the sign bit provided by the XOR array, and spliced ​​with the sign bit to obtain the results of each fixed-point multiplication; if the multiplier and the multiplicand are unsigned fixed-point types, the valid bits provided by the normalization array are directly used as the result; finally, the output post-processing module performs data splicing on all multiplication results in sequence to obtain the multiplication result

[0029] The reconfigurable multiplication-accumulation unit for a neural network processor, wherein the second input preprocessing module is used to: split the slice result according to the data type: if the data type to be added is a floating-point type, the exponent bit of each data pair is input into the exponent comparison array, and the sign bit and the mantissa bit are input into the shift array; if the data type to be added is an unsigned fixed-point type, the valid bit is directly input into the shift array; if the data type to be added is a signed fixed-point type, the sign bit and the valid bit are spliced ​​and then input into the shift array.

[0030] The reconfigurable multiplication-accumulation unit for a neural network processor, wherein the shift array includes multiple shifters for: when the input is floating-point data, the data is shifted according to the maximum exponent provided by the exponent comparison array and a sign bit is added to obtain a shift result; when the input is fixed-point data, the shift array does not work; and the shift result is finally passed to the addition array.

[0031] The reconfigurable multiplication-accumulation unit for a neural network processor, wherein the addition array is used to: if the data to be added is floating-point data, first perform a complement operation on the mantissa, then perform an addition operation, then convert the result back to the original code to obtain the sign bit and data bit, and finally pass the operation result to the normalized array; if the data to be added is fixed-point data, directly perform the addition operation, and pass the operation result to the normalized array.

[0032] The reconfigurable multiply-accumulate unit for a neural network processor, wherein the normalized array is used for:

[0033] When the data to be added is a floating-point number, the operation result of the addition array is received, and the exponent adjustment value is passed to the exponent adjustment array. The mantissa is shifted by the exponent adjustment value obtained above and the part exceeding the valid bit is truncated. The normalized mantissa bits are passed to the accumulation register. When the data to be added is a fixed-point data, the result is directly passed to the accumulation register.

[0034] The exponent adjustment value is calculated by using multiple sets of leading zero counters to count the number of leading zeros of each input element, and the counting result is the exponent adjustment value of each data element.

[0035] The reconfigurable multiplication-accumulation unit for a neural network processor, wherein the accumulation register is used for:

[0036] If the data to be added is floating-point type, the sign bit and mantissa bits provided by the normalization array and the exponent bits provided by the exponent adjustment array are concatenated to generate the result of the floating-point addition; if the data to be added is fixed-point type, the data provided by the normalization array is directly used; the above results are concatenated and cached in sequence to obtain the multiplication and addition result.

[0037] The present invention also proposes a neural network processor, which includes the reconfigurable multiplication and accumulation unit described in any one of claims 1-9.

[0038] From the above scheme, it can be seen that the advantages of the present invention are:

[0039] The design of the present invention supports various data types at 4, 8, 16, 32, and 64 bits within a single architecture, and simultaneously supports multiplication, multiply-add, and multiply-accumulate operations. Furthermore, independent multiply-accumulate modes and vector inner product modes are supported at low precision. As a result, the design of the present invention offers high flexibility and computational efficiency for neural network training and inference tasks, while effectively covering a wide range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 The figure is a diagram showing the overall structure of a reconfigurable multiplication-accumulation unit for a neural network processor according to the present invention;

[0041] Figure 2 The figure is a schematic diagram of the reconfigurable multiplier supporting multiplication of 64-bit data width;

[0042] Figure 3 The diagram shows the reconfigurable multiplier's support for 32-bit data width multiplication.

[0043] Figure 4 The diagram shows the reconfigurable multiplier's support for 16-bit data width multiplication.

[0044] Figure 5 The diagram shows the reconfigurable multiplier's support for 8-bit data width multiplication.

[0045] Figure 6 The diagram shows the reconfigurable multiplier's support for 4-bit data width multiplication. DETAILED DESCRIPTION

[0046] It should be noted that, in this application, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0047] Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.

[0048] The processor described in the present invention is the control center of an electronic device and can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).

[0049] Optionally, the processor can perform various functions of the electronic device by running or executing a software program stored in the memory, and calling data stored in the memory.

[0050] In a specific implementation, as an embodiment, the processor may include one or more CPUs. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). Electronic devices may include: servers, desktop computers, laptops, smartphones, tablet computers, embedded computers, etc., wherein the embedded computers include vehicles and robots, etc.

[0051] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0052] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto, and the actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or a combination of certain components, or a different arrangement of components.

[0053] The above embodiments can be implemented in whole or in part through software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired method (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0054] It should also be understood that the term "and / or" in this document simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " in this document generally indicates an "or" relationship between the related objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0055] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0056] It should also be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0057] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0058] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0059] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0060] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0061] This paper systematically analyzes the data types used in neural network algorithms to derive a set of commonly used data types. Research has shown that neural network computations require a high degree of diversity in the types of data they support, including high-precision and low-precision, floating-point and fixed-point. Designing a hardware architecture that efficiently accommodates these fragmented computational requirements presents a significant challenge.

[0062] To solve this problem, the inventors conducted an in-depth analysis of the computational behavior of these data types and proposed an optimized multiplier fusion solution. Specifically, by rationally organizing a large number of low-precision multiplication units, an efficient reconfigurable multiplier was constructed. Furthermore, the present invention designs a peripheral reconfigurable computing circuit to form a multiplication-accumulation unit, so that the unit can flexibly support the computational requirements of multiple precisions, multiple calculation types, and multiple data paths, thereby significantly improving the hardware's compatibility and efficiency with neural networks. In order to achieve the above technical effects, the present invention proposes the following key technical points:

[0063] Key point 1: A reconfigurable multiplier design that supports multiple calculation precisions; the technical effect is a reconfigurable multiplier that supports: 16 4-bit fixed-point multiplication operations; 8 8-bit fixed-point or floating-point mantissa multiplication operations; 4 16-bit fixed-point or floating-point mantissa multiplication operations; 2 32-bit fixed-point or floating-point mantissa multiplication operations; and 1 64-bit fixed-point or floating-point mantissa multiplication operation;

[0064] Key Point 2: A design and calculation method for a reconfigurable multiplication-accumulation unit for a neural network processor. The technical effect is that, in conjunction with the reconfigurable multiplier, it can be reconfigured to perform multiplication, multiplication-addition, and multiplication-accumulation operations at various precisions. Furthermore, when performing low-precision operations, the unit can be reconfigured into multiple independent multiplication-accumulators to implement outer product solutions for matrix multiplication, or into a single vector inner product unit to implement inner product solutions for matrix multiplication.

[0065] To illustrate the above-mentioned features and effects of the present invention more clearly and easily, the following embodiments are specifically described below with reference to the accompanying drawings. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are for illustrative purposes only. The scope of protection of the present invention is not limited to the disclosed embodiments; the present invention is defined by the appended claims.

[0066] Figure 1 The overall structure of a reconfigurable multiplication-accumulation unit for a neural network processor proposed by the present invention is shown. Figure 1 The circuit shown in FIG4 can complete the calculation process of multiplying A and B and adding the multiplication result to C. Specifically, it includes a multiplier, an accumulator, and an output selector.

[0067] The multiplier includes the following modules: input pre-processing module, XOR array, addition array, reconfigurable multiplier, normalization array, exponent adjustment array and output post-processing module.

[0068] The multiplier's input preprocessing module primarily performs data splitting and conversion. First, the module slices the two 64-bit input multiplicands and multipliers based on data precision. Specifically, if the multiplier and multiplicand have the same precision (x bits for both precisions, x = 4, 8, 16, 32, 64), the multiplicand and multiplier are sliced ​​into segments of x bits each. After slicing, the corresponding segments form a pair of data, totaling 64 / x pairs. Next, the module further splits or converts the input data based on data type. Specifically, if the input data pairs are floating-point, the sign bit of each pair is input into an exclusive-or array, the exponent bit of each pair is input into an adder array, and the mantissa bit of each pair is input into a reconfigurable multiplier. If the input data pairs are unsigned fixed-point, the significand bits of the data are directly input into the reconfigurable multiplier. If the input data pairs are signed fixed-point, the sign bit is input into an exclusive-or array, and the data's complement is taken and then input into the reconfigurable multiplier.

[0069] The XOR array primarily handles the sign bit calculation of multiplication results and consists of multiple XOR units. This module receives the sign bits of each data pair output by the input preprocessing module and performs XOR processing on them to obtain the sign bits of each multiplied data pair. The result is then passed to the output postprocessing module. This module does not function if processing unsigned fixed-point data.

[0070] The adder array primarily handles exponent calculations in floating-point operations and consists of multiple adders. It receives the exponent bits of each data pair output by the input preprocessing module and adds them together. The resulting exponent addition is then passed to the multiplier's exponent adjustment array. This module can be reconfigured or configured with a set of adders for each precision, performing exponent calculations on each pair of input data. This module is inactive when processing fixed-point data.

[0071] The reconfigurable multiplier is primarily used for fixed-point multiplication and floating-point mantissa multiplication. It contains one 32-bit x 32-bit multiplier, five 16-bit x 16-bit multipliers, thirteen 8-bit x 8-bit multipliers, and a set of adders. These multipliers can be flexibly combined to form a single high-precision multiplier or multiple independent low-precision multipliers.

[0072] The normalization array is primarily used to normalize the mantissa bits in floating-point operations and includes multiple shifters. This module receives the multiplication results of the reconfigurable multiplier and passes the exponent adjustment value to the exponent adjustment module. The exponent adjustment value is used to shift the mantissa and truncate the portion exceeding the valid bit to complete the normalization process. The normalized mantissa bits are then passed to the output post-processing module. If fixed-point data is being processed, the module directly passes the result to the output post-processing module without performing normalization. The exponent adjustment value is calculated by using multiple leading zero counters to count the number of leading zeros for each input element (in low-precision calculations, the input is composed of multiple elements spliced ​​together). The count result is the exponent adjustment value for each data element.

[0073] The exponent adjustment module (exponent adjustment array) is primarily used to further adjust the exponent of floating-point data. It consists of a set of adders or a set of data selectors. This module receives the exponent adjustment signal from the normalization array and adjusts the sum of the exponents of each data pair provided by the adder array, ultimately passing the result to the output post-processing module. This module does not function when processing fixed-point data.

[0074] The output post-processing module is primarily responsible for data merging and conversion. Specifically, if the input data is floating-point, the sign bit provided by the XOR array, the exponent bit provided by the exponent adjustment array, and the mantissa bits provided by the normalization array are concatenated to generate the results of each floating-point multiplication. If the input data is signed fixed-point, the significand bits provided by the normalization array are complemented by the sign bit provided by the XOR array and concatenated with the sign bit to obtain the results of each fixed-point multiplication. If the input data is unsigned fixed-point, the significand bits provided by the normalization array are directly used as the result. Finally, this module concatenates all the multiplication results in order.

[0075] The accumulator includes the following modules: input pre-processing module, exponent comparison array, shift array, addition array, normalization array, exponent adjustment array, accumulation register and data selector.

[0076] The accumulator input preprocessing module primarily performs data segmentation. First, it slices the two input data sets (i.e., the multiplier result and the accumulator register result) logarithmically. This is because the two accumulator inputs may consist of multiple concatenated pairs of data (for low-precision calculations). These concatenated data must be segmented and the corresponding pairs added. The segmentation is performed based on the actual data width of each input element. For example, if the width of a single data element in the multiplier result (one of the inputs) is 8 bits (4 bits x 4 bits), the multiplier result is segmented into 8-bit groups. The module then further segments the input data based on the data type of the addend. If the input data pairs are floating-point, the exponent bits of each pair are entered into the exponent comparison array, and the sign and mantissa bits are entered into the shift array. If the input data pairs are unsigned fixed-point, the significand bits are directly entered into the shift array. If the input data pairs are signed fixed-point, the sign and significand bits are concatenated and entered into the shift array.

[0077] The exponent comparison array, primarily used for comparing floating-point data exponents, contains multiple comparators. It accepts the exponent values ​​(exponent bits) of floating-point data provided by the input preprocessing module for comparison and passes the resulting maximum exponent to the shift array and exponent adjustment array. Specifically, if the multiply-accumulate unit is configured in independent multiply-accumulate mode, it compares each pair of data during low-precision calculations, outputting the maximum exponent for each pair. If the multiply-accumulate unit is configured in vector inner product mode, it compares multiple pairs of data as a whole during low-precision calculations, outputting the maximum exponent. This module does not function when processing fixed-point data.

[0078] The shift array, consisting of multiple shifters, is primarily used for data shift operations. It accepts the sign and mantissa bits of floating-point data, or the sign and significant bits of fixed-point data, provided by the input preprocessing module. When the input is floating-point data, the shift array shifts the data according to the maximum exponent provided by the exponent comparison array, thereby performing an exponential operation and adding the sign bit. When the input is fixed-point data, the shift array does not operate. The shifted result is ultimately passed to the addition array.

[0079] The adder array, primarily used to perform addition operations on the mantissas of fixed-point and signed floating-point data, contains multiple adders. It receives data from the shift array and processes it as follows: for floating-point data, the mantissa is first complemented, followed by addition. The result is then converted back to the original code to obtain the sign and data bits, and finally passed to the normalized array. For fixed-point data, the addition operation is performed directly, and the result is passed to the normalized array. In the specific implementation of low-precision addition: if the multiplication and accumulation unit is configured in independent multiplication and accumulation mode, the data within the data pair are added pair by pair, and a set of addition results are output; if the multiplication and accumulation unit is configured in vector inner product mode, multiple pairs of data are added together, and a sum is output.

[0080] The normalization array of the accumulator is mainly used for normalizing the mantissa in floating-point operations and includes multiple shifters. It receives the calculation results of the addition array and passes the exponent adjustment value to the exponent adjustment array. It shifts the mantissa using the exponent adjustment value obtained above and truncates the part exceeding the valid bit, and passes the normalized mantissa bits to the accumulator register. When processing fixed-point data, this module does not work, but directly passes the result to the accumulator register. The calculation method of the exponent adjustment value is similar to the above, using multiple sets of leading zero counters to count the number of leading zeros of each input element (in low-precision calculations, the input is spliced ​​from multiple elements). The counting result is the exponent adjustment value of each data element.

[0081] The accumulator's exponent adjustment array is primarily used to further adjust the exponent of floating-point data. It consists of a set of adders. This module receives the exponent adjustment value from the normalization array and adjusts it to the maximum exponent value provided by the exponent comparison array, ultimately passing the result to the accumulator register. This module is inactive when processing fixed-point data.

[0082] The accumulator register is primarily used for data caching and concatenation. Specifically, if the input data is floating-point, the sign and mantissa bits provided by the normalization array and the exponent bits provided by the exponent adjustment array are concatenated to generate the floating-point addition result. If the input data is fixed-point, the data provided by the normalization array is used directly without additional processing. Finally, the module concatenates and caches these results in sequence.

[0083] The two data selectors in the block diagram can reconfigure the multiplication-accumulation unit into three types at the computation level: multiplier, multiplication-accumulation unit, or multiplication-accumulation unit. Specifically, when selector m0 = S1, the output result D is the multiplication result; when selectors m0 = S0 and m1 = S1, the output result D is the multiplication-addition result; and when selectors m0 = S0 and m1 = S0, the output result D is the multiplication-accumulation result.

[0084] Table 1 lists the computational precisions supported by the reconfigurable multiply-accumulate unit, the number of equivalent computational components for the corresponding precision, supported data types, and the output bit widths of the multiplier and multiply-accumulator. This includes 13 data types, including the five fixed-point types (signed and unsigned) and eight floating-point types listed in Table 1. Supported computational precisions cover 4 bits, 8 bits, 16 bits, 32 bits, and 64 bits, meeting the computational requirements of virtually all quantized neural networks.

[0085] Accuracy Equivalent number Fixed-point type floating-point types Multiplication output bit width Multiply-accumulate output bit width 4bit 16 Int4 / Uint4 - 8bit 32bit 8bit 8 Int8 / Uint8 E3M4 / E4M3 / E5M2 16bit 32bit 16bit 4 Int16 / Uint16 FP16(E5M10) / BF16(E8M7) 32bit 32bit 32bit 2 Int32 / Uint32 FP32(E8M23) / TF32(E8M10) 64bit 64bit 64bit 1 Int64 / Uint64 FP64(E11M52) 64bit 64bit

[0086] Table 1

[0087] Figure 2 The reconfigurable multiplier supports 64-bit data width multiplication: the left image shows the reconfiguration into an Int64 / Uint64; the right image shows the reconfiguration into an FP64 (E11M52). The Uint64 reconfiguration uses the following hardware implementation: one 32-bit × 32-bit multiplier, four 16-bit × 16-bit multipliers, six 8-bit × 8-bit multipliers, and one adder group. The specific allocation is as follows:

[0088] 32-bit × 32-bit multiplier: Calculates the multiplication of inputs A[31:0] and B[31:0].

[0089] 16-bit × 16-bit multiplier: The first one is responsible for calculating A[15:0] × B[47:32]; the second one is responsible for calculating A[47:32] × B[15:0]; the third one is responsible for calculating A[31:16] × B[47:32]; and the fourth one is responsible for calculating A[47:32] × B[31:16].

[0090] 8-bit × 8-bit multiplier: The first one is responsible for calculating A[7:0] × B[55:48]; the second one is responsible for calculating A[55:48] × B[7:0]; the third one is responsible for calculating A[15:8] × B[55:48]; the fourth one is responsible for calculating A[55:48] × B[15:8]; the fifth one is responsible for calculating A[7:0] × B[63:56]; the sixth one is responsible for calculating A[63:56] × B[7:0].

[0091] Finally, the adder group shifts the results of each multiplier left and adds them together to obtain the Uint64 multiplication result, where: the 32-bit × 32-bit result does not need to be shifted; the first and second 16-bit × 16-bit results are shifted left by 32 bits; the third and fourth 16-bit × 16-bit and the first and second 8-bit × 8-bit results are shifted left by 48 bits; the third, fourth, fifth, and sixth 8-bit × 8-bit results are shifted left by 56 bits.

[0092] The calculation process of Int64 is similar to that of Uint64, but the effective data bit width is 63 bits (the sign bit

[63] does not participate in the multiplication operation); for FP64, all multipliers participate in the operation, but only for the mantissa bits ([52:0]); the exponent bit and sign bit do not participate in the calculation.

[0093] Figure 3 The reconfigurable multiplier demonstrates support for 32-bit data width multiplication: the left image shows the reconfiguration into two Int32 / Uint32s; the middle image shows the reconfiguration into two FP32s (E8M23); and the right image shows the reconfiguration into two TF32s (E8M10). A and B are data sets, each containing two 32-bit data elements, concatenated in the order shown. Both Int32 / Uint32 and FP32 are implemented using one 32-bit × 32-bit and four 16-bit × 16-bit multipliers and a set of adders; TF32, on the other hand, requires only two 16-bit × 16-bit multipliers.

[0094] Figure 4 The reconfigurable multiplier demonstrates support for 16-bit data width multiplication: the left image shows reconstruction into four Int16 / Uint16 arrays; the middle image shows reconstruction into four FP16 arrays (E5M10); and the right image shows reconstruction into four BF16 arrays (E8M7). A and B are each a set of four 16-bit data elements, concatenated in the order shown. These multiplications are performed by four 16-bit x 16-bit multipliers.

[0095] Figure 5 This figure demonstrates the reconfigurable multiplier's support for 8-bit data width multiplication: the upper left image shows reconstruction into eight Int8 / Uint8 arrays; the upper right image shows reconstruction into eight E3M4 floating-point arrays; the lower left image shows reconstruction into eight E4M3 floating-point arrays; and the lower right image shows reconstruction into eight E5M2 floating-point arrays. A and B are data sets, each containing eight 8-bit data elements, concatenated in the order shown. The calculations are performed using eight 8-bit x 8-bit multipliers.

[0096] Figure 6 This diagram demonstrates the reconfigurable multiplier's support for 4-bit data width multiplication, reconfiguring it as 16 Int4 / Uint4 operations. A and B each contain 16 4-bit data elements, concatenated in the order shown. These calculations are performed by 13 8-bit × 8-bit and three 16-bit × 16-bit multipliers.

[0097] It should be noted that Figures 2 to 6 This is only one example of the reconfigurable multiplier proposed in the present invention. In practical applications, the mapping relationship between calculation and hardware, and the size and number of multipliers can be flexibly adjusted according to needs.

Claims

1. A reconfigurable multiply-accumulate unit for a neural network processor, characterized in that: include: Multiplier, accumulator; The multiplier includes: a first input pre-processing module, an XOR array, an addition array, a reconfigurable multiplier, a normalization array, an exponent adjustment array and an output post-processing module; The first input preprocessing module is configured to slice the multiplier and the multiplicand according to their data precision to obtain a plurality of data pairs; input the sign bit of each data pair into the XOR array, input the exponent bit of each data pair into the addition array, and input the mantissa bit or the valid bit of each data pair into the reconfigurable multiplier according to the data type of the data pair; The XOR array includes a plurality of XOR units for receiving the sign bits of each data pair, performing XOR processing on the sign bits, and obtaining the sign bits of the multiplied data pairs; The addition array includes a plurality of adders for receiving the exponent bits of each data pair, performing addition processing on the exponent bits, obtaining the exponent addition result of each data pair, and passing the result to the exponent adjustment array; The reconfigurable multiplier includes a plurality of multipliers and a group of adders for performing fixed-point multiplication and floating-point mantissa multiplication operations to obtain a multiplication result; The normalization array includes a plurality of shifters for receiving the multiplication result and passing the exponent adjustment value to the exponent adjustment module, and passing the normalized mantissa bits to the output post-processing module; The exponent adjustment array includes an adder or a data selector for receiving the exponent adjustment value to adjust the exponential addition result and pass the adjustment result to the output post-processing module; The output post-processing module is used to obtain a multiplication result according to the sign bit, the adjustment result and the mantissa bit; The accumulator includes: a second input preprocessing module, an exponent comparison array, a shift array, an addition array, a normalization array, an exponent adjustment array, and an accumulation register; The second input preprocessing module slices the logarithms of the data to be added, splits the slicing results according to data types, inputs the exponent bits obtained by the splitting into the exponent comparison array, inputs the sign bit and the mantissa bits into the shift array, and splices the significant bits or the sign bit and the significant bits into the shift array; The exponent comparison array comprises a plurality of comparators for comparing the exponent bits and passing the maximum exponent bit to the shift array and the exponent adjustment array; The shift array includes a plurality of shifters, which are used to receive data provided by the second input pre-processing module and shift the data according to the maximum exponent bit to obtain a processing result and pass it to the addition array; The addition array includes a plurality of adders, which are used to perform addition operations on the mantissas of the fixed-point data and the signed floating-point data according to the processing results to obtain operation results. The normalization array includes a plurality of shifters for receiving the operation result, passing the exponent adjustment value to the exponent adjustment array, and passing the normalized mantissa bits in the operation result to the accumulation register; The exponent adjustment array comprises an adder for adjusting the maximum exponent bit according to the exponent adjustment value and passing the adjustment result to the accumulation register; The accumulator register is used to cache and merge data to obtain the multiplication and addition result.

2. The reconfigurable multiplication-accumulation unit for a neural network processor according to claim 1, wherein: Also includes: Output selector; the input of the output selector is the output of the accumulator and the output of the multiplier, and is used to select the output of the accumulator or the output of the multiplier according to the control instruction; The accumulator also includes a data selector; the input of the data selector is the multiplication and addition result output by the accumulator register and the addend input to the accumulator, and is used to select, according to the control instruction, whether the data input to the second input preprocessing module is the multiplication and addition result output by the accumulator register or the addend.

3. The reconfigurable multiplication-accumulation unit for a neural network processor according to claim 1, wherein: The first input preprocessing module is used to: If the multiplier and the multiplicand have the same precision of x bits, the multiplicand and the multiplier are sliced ​​into segments of x bits each, and the corresponding segments after slicing constitute a pair of data, with a total of 64 / x pairs; if the input data pairs are floating-point type, the sign bit of each data pair is input into the XOR array, the exponent bit of each data pair is input into the addition array, and the mantissa bit of each data pair is input into the reconfigurable multiplier; if the input data pairs are unsigned fixed-point type, the valid bits of the data are directly input into the reconfigurable multiplier; if the input data pairs are signed fixed-point type, the sign bit is input into the XOR array, and the data complement is taken from the original code and then input into the reconfigurable multiplier.

4. The reconfigurable multiplication-accumulation unit for a neural network processor according to claim 1, wherein: The output post-processing module is responsible for data merging and conversion. If the multiplier and the multiplicand are floating-point types, the sign bit provided by the XOR array, the exponent bit provided by the exponent adjustment array, and the mantissa bit provided by the normalization array are spliced ​​to generate the results of each floating-point multiplication; if the multiplier and the multiplicand are signed fixed-point types, the valid bits provided by the normalization array are complemented according to the sign bit provided by the XOR array and spliced ​​with the sign bit to obtain the results of each fixed-point multiplication; if the multiplier and the multiplicand are unsigned fixed-point types, the valid bits provided by the normalization array are directly used as the result; finally, the output post-processing module performs data splicing on all multiplication results in sequence to obtain the multiplication result.

5. The reconfigurable multiplication-accumulation unit for a neural network processor according to claim 1, wherein: The second input preprocessing module is used to: split the slice result according to the data type: if the data type to be added is floating point type, input the exponent bit of each data pair into the exponent comparison array, and input the sign bit and the mantissa bit into the shift array; If the data type to be added is an unsigned fixed-point type, the valid bits are directly input into the shift array; If the data type to be added is a signed fixed-point type, the sign bit and the valid bit are concatenated and then input into the shift array.

6. The reconfigurable multiplication-accumulation unit for a neural network processor according to claim 1, wherein: The shift array includes multiple shifters for: when the input is floating-point data, the data is shifted according to the maximum exponent provided by the exponent comparison array and a sign bit is added to obtain a shift result; when the input is fixed-point data, the shift array does not work; and the shift result is finally passed to the addition array.

7. The reconfigurable multiplication-accumulation unit for a neural network processor according to claim 1, wherein: The addition array is used to: if the data to be added is floating-point data, first perform a two's complement operation on the mantissa, then perform an addition operation, then convert the result back to the original code to obtain the sign bit and data bits, and finally pass the operation result to the normalized array; If the data to be added is fixed-point data, the addition operation is performed directly, and the operation result is passed to the normalized array.

8. The reconfigurable multiplication-accumulation unit for a neural network processor according to claim 1, wherein: This normalized array is used to: When the data to be added is a floating-point number, the operation result of the addition array is received, and the exponent adjustment value is passed to the exponent adjustment array. The mantissa is shifted by the exponent adjustment value obtained above and the part exceeding the valid bit is truncated. The normalized mantissa bits are passed to the accumulation register. When the data to be added is a fixed-point data, the result is directly passed to the accumulation register. The exponent adjustment value is calculated by using multiple sets of leading zero counters to count the number of leading zeros of each input element, and the counting result is the exponent adjustment value of each data element.

9. The reconfigurable multiplication-accumulation unit for a neural network processor according to claim 1, wherein: This accumulator register is used to: If the data to be added is of floating-point type, the sign bit and mantissa bits provided by the normalization array and the exponent bits provided by the exponent adjustment array are concatenated to generate the result of the floating-point addition; If the data to be added is of fixed-point type, the data provided by the normalized array is directly used; the above results are sequentially concatenated and cached to obtain a multiplication-addition result.

10. A neural network processor, characterized in that: The invention comprises the reconfigurable multiplication-accumulation unit described in any one of claims 1-9.

Citation Information

Cited By

  • Error upper bound device suitable for dynamic precision floating point multiply-accumulate operation, judgment method, medium, terminal and program product

    CN121523639A