A multiple precision parallel multiplication array structure and method

CN122526534BActive Publication Date: 2026-09-08XIAN UNIV OF POSTS & TELECOMM +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611015758.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-08
Estimated Expiration
2046-07-09

AI Technical Summary

Technical Problem

在资源利用方面,传统设计通常为不同位宽的乘法运算分别配置独立的乘法器模块,导致硬件资源出现大量冗余,芯片面积开销显著增加,资源利用率低;在模块复用方面,部分可重构设计虽实现了一定程度的资源复用,但复用粒度较粗,多以16bit块等较大模块为复用单位,缺乏细粒度的复用机制,灵活性不足,无法高效适配多种精度组合的运算需求;在结构扩展方面,现有阵列结构乘法器在扩展运算位宽时,需要对整体结构进行重新设计,缺乏统一的阵列映射机制,扩展性较差;在模式切换方面,不同精度模式之间的切换需要复杂的控制逻辑支撑,数据路径难以实现有效复用,导致切换效率低,影响整体运算性能

Benefits of technology

本发明提供了一种多精度并行乘法阵列结构,构建一个由多个相同基础乘法单元组成的二维阵列,每个单元作为最小复用单元执行预设位宽乘法,通过行级与列级组合灵活构造不同位宽乘法器;同时配备乘数分解与映射模块将操作数按位宽分解映射至对应单元,并由压缩求和模块对单元输出进行分层并行压缩及求和得到最终结果。本结构设计通过细粒度的基础单元复用机制,将大位宽乘法分解为多个小位宽并行计算,结合二维阵列的规则排布与统一映射,实现了硬件资源的动态重组与数据路径高效复用。采用本结构显著提升了资源利用率,避免了传统设计的冗余问题;增强了可重构性,能够适配多种精度组合运算;改善了结构扩展性,无需重新设计即可扩展位宽;同时优化了模式切换效率,简化控制逻辑,从而在高资源利用率、灵活可重构性与高性能之间实现有效平衡。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122526534B_ABST
    Figure CN122526534B_ABST
Patent Text Reader

Abstract

The application belongs to the field of multipliers, and discloses a multi-precision parallel multiplication array structure and method, which constructs a two-dimensional array composed of multiple same basic multiplication units, each unit serving as a minimum reuse unit to perform preset bit-width multiplication, and different bit-width multipliers are flexibly constructed through row-level and column-level combination; meanwhile, a multiplier decomposition and mapping module is provided to decompose and map the operands to corresponding units according to bit width, and a compression and summation module is provided to perform layered parallel compression and summation on the unit outputs to obtain the final result. The structure significantly improves the resource utilization, avoids the redundancy problem of traditional design, enhances the reconfigurability, can adapt to various precision combination operations, improves the structure scalability, can be expanded in bit width without redesign, optimizes the mode switching efficiency, simplifies the control logic, and thus effectively balances the high resource utilization, flexible reconfigurability and high performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multiplier technology, and particularly relates to a multi-precision parallel multiplication array structure and method. Background Technology

[0002] Multipliers are core computing units in graphics processing units (GPUs), artificial intelligence accelerators, and general-purpose processors. Their performance directly determines the overall efficiency and power consumption of the entire computing system, playing an irreplaceable role in fields such as artificial intelligence, scientific computing, and data processing that rely on high-speed computation. Currently, mainstream multiplier designs in the industry are mainly divided into three categories: Booth-based multipliers, high-speed multipliers based on array structures (such as Wallace trees and Dadda trees), and fixed-width multipliers (such as 8-bit, 16-bit, and 32-bit). In practical applications, different computing tasks have significantly different requirements for data precision. Low-precision (e.g., 8-bit) multiplication is widely used in scenarios such as neural network inference where precision requirements are lower but speed and energy efficiency are higher. High-precision (e.g., 16-bit and 32-bit) multiplication is mainly used in scenarios such as scientific computing where data precision requirements are strict. Based on this, in recent years, "multi-precision reconfigurable multipliers" have become a research hotspot in related fields due to their ability to adapt to diverse precision requirements.

[0003] Although research on multi-precision reconfigurable multipliers has made some progress, existing technologies still have many shortcomings and deficiencies. Regarding resource utilization, traditional designs typically configure separate multiplier modules for multiplication operations of different bit widths, leading to significant hardware redundancy, a substantial increase in chip area overhead, and low resource utilization. Regarding module reuse, while some reconfigurable designs achieve a certain degree of resource reuse, the reuse granularity is coarse, often using large modules such as 16-bit blocks as reuse units, lacking fine-grained reuse mechanisms, resulting in insufficient flexibility and an inability to efficiently adapt to the computational needs of various precision combinations. Regarding structural expansion, existing array-structured multipliers require a complete redesign of the entire structure when expanding the operational bit width, lacking a unified array mapping mechanism and exhibiting poor scalability. Regarding mode switching, switching between different precision modes requires complex control logic, making it difficult to effectively reuse data paths, resulting in low switching efficiency and impacting overall computational performance.

[0004] It is evident that existing multiprecision multiplier designs struggle to achieve an effective balance between high resource utilization, flexible reconfigurability, and high performance, failing to fully meet the diverse adaptation needs of various computing tasks for multipliers. Summary of the Invention

[0005] This invention provides a multi-precision parallel multiplication array structure and method. The multiplication array structure can achieve an effective balance between high resource utilization, flexible reconfigurability and high performance, and can fully meet the diverse adaptation needs of multipliers for various computing tasks.

[0006] To achieve the above objectives, the present invention employs the following technical content: A multi-precision parallel multiplication array structure includes: A two-dimensional multiplication array structure; the two-dimensional multiplication array structure includes multiple basic multiplication units with the same structure; the multiple basic multiplication units are arranged in a two-dimensional array and constructed by row-level and column-level combinations to form multipliers with different bit widths; each of the basic multiplication units serves as a minimum multiplexing unit for performing multiplication operations with a preset bit width; The multiplier decomposition and mapping module is used to receive operands of the target precision, decompose the operands into multiple data segments according to the bit width of the operands, and map the multiple data segments to the corresponding basic multiplication units in the two-dimensional multiplication array structure. The compression and summation module is used to perform hierarchical parallel compression and summation on the operation results output by the basic multiplication unit to obtain the multiplication result with the target precision.

[0007] Furthermore, the basic multiplication unit is also used to switch between signed number multiplication operation mode and unsigned number multiplication operation mode based on the received configuration signal.

[0008] Furthermore, the basic multiplication unit includes: The encoding subunit is used to encode operands of a preset bit width using the radix-4 Booth encoding algorithm to generate a partial product; The symbol processing subunit specifically includes: a first variable generation circuit for generating a first variable for processing the sign bit extension; and a second variable generation circuit for generating a second variable for uniformly processing the partial product sign in both signed and unsigned multiplication operation modes.

[0009] Furthermore, the row direction of the two-dimensional multiplication array structure is configured to support parallel computation of first-precision data, and the column direction of the two-dimensional array is configured to support parallel computation of second-precision data; wherein the bit width of the first-precision data is lower than the bit width of the second-precision data.

[0010] Furthermore, the basic multiplication unit adopts an 8-bit sub-multiplier or a 4-bit sub-multiplier; the two-dimensional multiplication array structure adopts a regular array of 9 rows and 4 columns, integrating a total of 36 basic multiplication units.

[0011] Furthermore, the compression and summation module includes a multi-level compression architecture, configured to compress and normalize the irregularly distributed operation results output by the basic multiplication unit into two rows of output data through hierarchical parallel compression, and add the two rows of output data to obtain the multiplication result with the target precision.

[0012] Furthermore, the multi-stage compression architecture includes one or more of a 4-2 compressor, a 3-2 compressor, and a half-compressor; wherein: The 4-2 compressor is configured to compress four equally weighted input bits and one low-carry input into one sum bit and two carry outputs; The 3-2 compressor is configured to compress three equally weighted input bits into one output bit and one carry; The semi-compressor is configured to compress two inputs into one output and one carry.

[0013] Furthermore, the execution logic expression of the 4-2 compressor is as follows:

[0014]

[0015]

[0016] The execution logic expression of the 3-2 compressor is as follows:

[0017]

[0018] The execution logic expression of the semi-compressor is as follows:

[0019]

[0020] In the formula, , , , These are the compressor input bits; This is a carry-in from the lower digit. For compressor and bit output; Outputting the carry from the middle; For carry-out output; It is the XOR operator.

[0021] Furthermore, when the target precision is the first target precision, the multiplier decomposition and mapping module is configured to decompose the operand into high-bit data segments, middle-bit data segments and low-bit data segments according to the bit width of the basic multiplication unit, and map them to multiple basic multiplication units in the same column; When the target precision is the second target precision, the multiplier decomposition and mapping module is configured to decompose the operand into high-bit data segments and low-bit data segments according to the bit width of the basic multiplication unit, and map them to multiple basic multiplication units in the same row. The bit width of the first target precision is higher than that of the second target precision.

[0022] A multi-precision parallel multiplication operation method, based on the aforementioned multi-precision parallel multiplication array structure, includes: The multiplier decomposition and mapping module receives operands of the target precision, decomposes the operands into multiple data segments according to the bit width of the operands, and maps the multiple data segments to the corresponding basic multiplication units in the two-dimensional multiplication array structure. The basic multiplication unit, employing a two-dimensional multiplication array structure, performs multiplication operations of a preset bit width on multiple data segments and outputs the results. A compression and summation module is used to perform hierarchical parallel compression and summation on the operation results output by the basic multiplication unit to obtain the multiplication result with the target precision.

[0023] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a multi-precision parallel multiplication array structure, constructing a two-dimensional array composed of multiple identical basic multiplication units. Each unit serves as the minimum reuse unit, performing multiplication with a preset bit width. Different bit width multipliers are flexibly constructed through row-level and column-level combinations. Simultaneously, a multiplier decomposition and mapping module decomposes and maps operands to corresponding units according to their bit width, and a compression and summation module performs hierarchical parallel compression and summation on the unit outputs to obtain the final result. This structure design, through a fine-grained basic unit reuse mechanism, decomposes large-bit-width multiplication into multiple small-bit-width parallel computations. Combined with the regular arrangement and unified mapping of the two-dimensional array, it achieves dynamic reorganization of hardware resources and efficient reuse of data paths. This structure significantly improves resource utilization, avoids redundancy problems in traditional designs, enhances reconfigurability, adapts to various precision combinations, improves structural scalability (bit width can be expanded without redesign), and optimizes mode switching efficiency and simplifies control logic, thus achieving an effective balance between high resource utilization, flexible reconfigurability, and high performance.

[0024] This invention also provides a multi-precision parallel multiplication method. Based on the aforementioned multi-precision parallel multiplication array structure, this method first decomposes the operands of the target precision into multiple data segments according to their bit width and maps them to the corresponding basic multiplication units in the array. These units then perform multiplication operations of a preset bit width in parallel. Finally, all the resulting partial product results are subjected to hierarchical parallel compression and summation to obtain the final product. By decomposing large-bit-width multiplication operations into multiple small-bit-width, regularized sub-operation tasks, and through fine-grained basic unit reuse and a unified array mapping mechanism, dynamic allocation and efficient parallel processing of computational tasks on hardware resources are achieved. This method fully utilizes the flexibility of the underlying array structure, significantly improving hardware resource utilization and avoiding redundancy; through fine-grained decomposition and mapping, it supports operations with multiple precision combinations, enhancing the system's flexibility and reconfigurability; the regularized operation process and unified path simplify the switching control between different precision modes, improving computational efficiency and overall performance, thus achieving an effective balance between resource utilization, reconfigurability, and high performance. Attached Figure Description

[0025] Figure 1 This is a schematic diagram of a multi-precision parallel multiplication array structure provided in an embodiment of the present invention; Figure 2 The partial product optimization process provided in this embodiment of the invention includes: (a) a schematic diagram of all positive partial products of unsigned 8-bit multiplication; (b) a schematic diagram of all negative partial products of unsigned 8-bit multiplication; (c) a schematic diagram of optimization of all negative partial products of unsigned 8-bit multiplication; (d) a schematic diagram of normal partial products of unsigned 8-bit multiplication; and (e) a schematic diagram of partial products of signed 8-bit multiplication. Figure 3 A schematic diagram of a partial product compression circuit supporting signed and unsigned 8-bit multiplication provided in an embodiment of the present invention; Figure 4 The following are schematic diagrams of compressors provided for embodiments of the present invention; wherein, (a) is a schematic diagram of a half compressor; (b) is a schematic diagram of a 3-2 compressor; and (c) is a schematic diagram of a 4-2 compressor. Figure 5 A schematic diagram of multiplier reconstruction provided for an embodiment of the present invention; Figure 6 This is a structural diagram of a 16-bit multiplication compression summation module provided in an embodiment of the present invention; Figure 7 This is a structural diagram of a 24-bit multiplication compression summation module provided in an embodiment of the present invention; Figure 8 A data allocation mapping provided in an embodiment of the present invention; Figure 9 Another data allocation mapping provided in this embodiment of the invention; Figure 10 This is a schematic diagram of unpacking provided in an embodiment of the present invention. Detailed Implementation

[0026] To make the technical problems solved by the present invention, the technical solutions, and the beneficial effects clearer, the following specific embodiments provide a further detailed description of the present invention. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of the invention.

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0028] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0029] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0030] The technical terms involved in this invention are explained as follows: GPU: Graphics Processing Unit.

[0031] Wallace tree: A tree structure for partial product compression in high-speed multipliers.

[0032] Dadda Tree: A classic partial product compression tree multiplier structure.

[0033] Booth encoding: an optimized encoding method for binary signed number multiplication. It re-encodes the multiplier, reduces the number of partial products generated by multiplication, simplifies the hardware circuit of the multiplier, and reduces the operation latency.

[0034] Radix-4 Booth encoding algorithm: also known as Radix-4 Booth, is a high-order Booth encoding algorithm with a radix of 4. It encodes every 2 bits of multiplier in groups, which directly reduces the partial product by half compared to radix-2 Booth encoding, greatly optimizing the area and delay of high-speed multipliers.

[0035] bit: A binary digit is the smallest unit of information measurement in a digital system.

[0036] FP32: 32-bit single-precision floating-point number.

[0037] FP16: Floating Point 16, which is a 16-bit half-precision floating point number.

[0038] BF16: short for Brain Floating Point 16, which is a floating-point number with brain precision.

[0039] INT8: Full name Integer 8-bit, which is an 8-bit integer.

[0040] UINT8: Full name Unsigned Integer 8-bit, which is an 8-bit unsigned integer.

[0041] As described in the background section, existing technologies typically employ a scheme of configuring independent multipliers for different bit widths to support calculations of various precisions. This results in severe hardware resource redundancy and large area overhead. Although some existing designs attempt to use reconfigurable designs to support reuse, they are often reconfigured on a large module basis (such as a 16-bit multiplier module), resulting in coarse-grained reuse, insufficient flexibility, and difficulty in efficiently adapting to multiple precision combinations. Furthermore, existing array multipliers usually require redesigning their internal structure when expanding their bit width, lacking a unified array mapping mechanism, leading to complex control logic during multi-precision switching and difficulty in reusing data paths.

[0042] To address the aforementioned issues, this embodiment provides a multi-precision parallel multiplication array structure. This array structure utilizes the two-dimensional array arrangement of basic multiplication units and the data decomposition mapping mechanism to achieve efficient reuse of hardware resources and flexible adaptation to multi-precision computation.

[0043] like Figure 1 As shown, this embodiment provides a multi-precision parallel multiplication array structure, including: a two-dimensional multiplication array structure, a multiplier decomposition and mapping module, and a compression summation module; wherein, the two-dimensional multiplication array structure is composed of multiple basic multiplication units with the same structure.

[0044] The basic multiplication unit is configured to perform multiplication operations with a preset bit width. Specifically, the basic multiplication unit is the smallest computational granularity constituting the entire array structure, and its computational bit width is preset to a fixed value. Those skilled in the art will understand that this preset bit width can be flexibly configured according to actual application scenarios; for example, it can be set to 8 bits, 4 bits, or other suitable bit widths, as long as it can serve as the basic unit for constructing higher bit-width operations. In this embodiment, the basic multiplication unit, as a standardized computational "building block," has a highly regular and consistent internal structure, enabling it to independently complete a low-precision multiplication operation. By combining and reusing these standardized basic units, the hardware resource redundancy caused by designing dedicated multipliers for each target precision is avoided, thereby significantly improving resource utilization.

[0045] A two-dimensional multiplication array structure is configured to arrange multiple basic multiplication units in a two-dimensional array. Specifically, this structure builds a physical pool of computing resources by arranging a large number of basic multiplication units according to a regular row and column orientation. This two-dimensional array structure provides the spatial basis for subsequent data mapping. Figure 1 In the architecture shown, the basic multiplication units are distributed in a grid pattern in space. This regular arrangement not only facilitates the physical implementation of the chip layout and reduces wiring complexity, but more importantly, it provides the possibility of parallel processing for data of different precisions. For example, different rows or columns in the array can be dynamically allocated to different computing tasks, thereby achieving flexible scheduling and space reuse of computing resources at the hardware level.

[0046] The multiplier decomposition and mapping module is configured to receive operands of the target precision, decompose the operands into multiple data segments, and map these data segments to corresponding basic multiplication units in a two-dimensional array. Specifically, the multiplier decomposition and mapping module is the core logic control unit for achieving multi-precision adaptation. When the input operand's bit width is greater than the preset bit width of the basic multiplication unit, the multiplier decomposition and mapping module can, according to preset decomposition rules, divide the high-bit-width operand into several data segments matching the preset bit width. For example, if the target precision is 24 bits and the basic unit is 8 bits, this module can decompose the 24-bit operand into three 8-bit data segments: high, medium, and low. Subsequently, the multiplier decomposition and mapping module maps these data segments to specific basic multiplication units in the array, enabling these units to work in parallel and collaboratively complete tasks that would otherwise require a high-bit-width multiplier. This "divide and conquer" strategy allows a single hardware array to be compatible with multiple different data precisions, greatly improving the system's versatility and flexibility.

[0047] The compression and summation module is configured to perform hierarchical parallel compression of the operation results output by the basic multiplication units to obtain the multiplication result with the target precision. Specifically, since the multiplier decomposition and mapping module breaks down a complete multiplication operation into multiple basic units that are executed in parallel, the output of each basic unit is only a partial product or intermediate result, which differs in weight and bit width. The role of the compression and summation module is to quickly aggregate these scattered and irregular intermediate results. It uses specific compression logic (such as carry-preserving adder trees) to compress multiple rows of partial products into two final rows of values, which are then used by the adder to obtain the final multiplication result. This module ensures that the decomposed calculation results can be correctly and efficiently restored to the final value with the target precision, guaranteeing the correctness of the calculation.

[0048] Based on the above scheme, this embodiment constructs a two-dimensional reconfigurable computing array based on fine-grained basic units. The basic multiplication units provide standardized computing capabilities, the two-dimensional multiplication array structure provides a spatial resource pool, the multiplier decomposition and mapping module realizes dynamic adaptation of tasks to resources, and the compression and summation module completes the final aggregation of results. This architecture design breaks the limitation of the fixed bit width of traditional multipliers, enabling the same set of hardware resources to be dynamically reorganized according to the precision requirements of the input data. While ensuring computing performance, it significantly reduces hardware overhead and solves the problems of low resource utilization and poor flexibility of existing multi-precision multipliers.

[0049] As another preferred embodiment of the present invention, this embodiment provides a detailed description of the internal structure and optimization mechanism of the basic multiplication unit. The preset bit width is a fixed value. Specifically, the preset bit width can be set to 8 bits, 4 bits, or other suitable values ​​according to actual application requirements. In a preferred embodiment, the preset bit width is set to 8 bits. 8 bits is chosen as the basic granularity because it has a natural advantage in processing common low-precision data (such as INT8), and at the same time, it strikes a good balance between the required array size and interconnect complexity when constructing higher-precision operations (such as 16-bit and 24-bit). If the granularity is too small (such as 2 bits), it will lead to a large array size and complex control logic; if the granularity is too large (such as 16 bits), the flexibility of adapting to low-precision data will be reduced, easily resulting in resource waste.

[0050] The basic multiplication unit also switches between signed and unsigned multiplication modes based on received configuration signals. Specifically, the basic multiplication unit integrates mode selection logic. When a configuration signal indicating signed multiplication is received, the unit processes the sign bit according to the two's complement multiplication rules; when a configuration signal indicating unsigned multiplication is received, the unit processes it according to the unsigned multiplication rules. This design allows the same physical circuit to be compatible with both INT8 (signed) and UINT8 (unsigned) data types without requiring separate multiplier hardware for each mode, thus further improving the array's versatility and resource utilization.

[0051] To achieve the above functions and optimize performance, the basic multiplication unit includes an encoding subunit and a symbol processing subunit.

[0052] The encoding subunit is configured to encode operands of a preset bit width using the Radix-4 Booth encoding algorithm to generate a partial product. Combined with... Figure 2 As shown, a traditional 8-bit multiplier performing array multiplication directly would generate 8 rows of partial products, resulting in excessively long subsequent summation paths and significant latency. This embodiment employs the Radix-4 Booth encoding algorithm, grouping the multipliers into three-bit blocks (with one overlapping bit). For 8-bit operands, after high-bit sign extension and low-bit zero padding, 5 partial products can be generated. Compared to traditional methods, the number of partial products is reduced by nearly half, significantly decreasing the logic depth of the subsequent compressed summation stage and thus greatly improving computational speed.

[0053] The sign processing subunit is configured to introduce a first variable and a second variable. The first variable is used to handle sign bit extension, and the second variable is used to uniformly handle the partial product sign under both the first and second operation modes. This is one of the core optimization points of this embodiment. During Booth encoding, the partial product may contain negative values, requiring sign bit extension. Directly extending the sign bit would generate a large number of consecutive "1"s, increasing not only the circuit switching power consumption but also the burden on the adder.

[0054] Specifically, the first variable (S variable) is used to handle sign bit extension. For negative partial products, traditional methods require filling all high-order bits with "1". This embodiment introduces the S variable, transforming the sign bit extension calculation into processing the S variable itself. For example, in... Figure 2 In the optimization process shown in (c) and (d), the carry effect caused by the sign bit extension is pre-calculated and represented by the S variable, which avoids directly filling the part product matrix with a long sign bit, thereby simplifying the circuit connection and reducing the capacitor load.

[0055] The second variable (E variable) is used to uniformly handle the sign of the partial product in both signed and unsigned modes. In signed multiplication, the multiplicand itself may be negative, and the partial product may also be negative, resulting in a "negative times negative equals positive" situation. In unsigned multiplication, only the sign of the partial product needs to be considered. To ensure compatibility between these two modes within the same circuit, this embodiment introduces the E variable. The logical value of the E variable is obtained by XORing the sign bit of the multiplicand with the sign bit generated by the Booth encoding. Through the E variable, the circuit can automatically identify whether the current operation is signed or unsigned and correct the highest bit of the partial product accordingly. This design cleverly eliminates the differences in circuit logic between signed and unsigned operations, allowing the basic multiplication unit to switch modes with a simple control signal without adding extra correction circuit overhead.

[0056] In addition, combined Figure 4 As shown, the partial product matrix optimized by the Booth algorithm exhibits irregular column distribution with varying bit counts in each column. Using a traditional regular array for column-by-column addition would increase the number of addition stages and lengthen the critical path. Therefore, this design employs a fast summation structure based on a compressor tree (compression summation module) to perform hierarchical parallel compression of the partial product. The compression summation module uses a compression tree structure (multi-level compression architecture) consisting of 4-2 compressors, 3-2 compressors, and half-compressors. Since the partial product matrix optimized by Booth encoding may exhibit irregular column distribution (e.g., some columns have 5 bits, some have 3 bits), the compression summation module dynamically selects the compressor type based on the actual height of each column. For example, columns with more bits are preferentially used with a 4-2 compressor to quickly reduce height, while columns with fewer bits use a 3-2 compressor or a half-compressor. Through multi-stage compression, all partial products are eventually normalized into two rows of output (usually Sum and Carry), and then the final adder is used to obtain a 16-bit multiplication result. This hierarchical compression strategy effectively balances circuit area and critical path delay, ensuring the high performance of the basic multiplication unit.

[0057] As can be seen, this embodiment reduces the number of partial products through Radix-4 Booth encoding and optimizes the sign processing logic through S and E variables, realizing hardware reuse of signed and unsigned operations. This not only improves the operational efficiency of a single basic multiplication unit, but also lays a solid foundation for the high-density integration and flexible reconfiguration of the entire array structure.

[0058] It should be noted that the 8-bit sub-multiplier, as the smallest multiplexing unit, takes two 8-bit operands as input and outputs a 16-bit multiplication result. Internally, it is implemented using radix-4 Booth encoding and a compressor tree (such as a Wallace tree). This module serves as the smallest multiplexing granularity unit.

[0059] Since this 8-bit sub-multiplier unit is instantiated extensively in the array, optimizing it can significantly improve the performance of the entire circuit. Furthermore, to adapt to more application scenarios, this sub-multiplier unit supports both signed and unsigned multiplication of 8-bit data, and the multiplier can be configured via control signals to perform the current signed / unsigned multiplication.

[0060] First, using the Radix-4 Booth encoding table shown in Table 1, the original eight partial products in the sub-multiplication unit are reduced to five. Furthermore, the S variable is introduced to handle the preprocessing of sign bit extension caused by the booth algorithm, and the E variable is introduced to unify the processing of signed and unsigned partial products. The specific optimization methods for partial products are as follows: Table 1 shows the partial product selection table for Radix-4 Booth.

[0061] This design uses the Radix-4 Booth encoding algorithm to re-encode the multipliers. The core idea is to encode every two bits of the multiplier to reduce the number of partial products. Using Radix-4 Booth encoding can significantly optimize the speed, area and power consumption of the multiplier while keeping the hardware complexity under control.

[0062] Radix-4 Booth encoding first expands the multiplier by bit width. The expansion method depends on the total bit width of the data. If the total bit width is even, two bits are added to the high-order bits and one bit to the low-order bits; if the total bit width is odd, one bit is added to the high-order bits and one bit to the low-order bits. This ensures an encoding pattern of every three bits (with the first bit overlapping by one bit). For 8-bit data, two bits are added to the high-order bits. For signed numbers, two sign bits are added; for unsigned numbers, 0 bits are added. The low-order bits are always added with one 0 bit. Radix-4 Booth encoding of the expanded 11-bit multiplier yields 5 corresponding partial products, a reduction of 3 from the original 8 partial products.

[0063] Introducing Radix-4 Booth encoding reduces the number of partial products, but the additional cost is the need to handle the inversion and addition of 1, shift operations, and high-bit sign extension calculations required when adding or subtracting the multiplicand by 2, 3, and 4. However, these aspects can be optimized.

[0064] In this embodiment, we first consider the case where the 8-bit data is an unsigned number. Assuming all partial products are positive, such as... Figure 2As shown in (a), ignoring the high-order padding of 0, although the data is all positive, the partial product width is 9 bits because there is a possibility of adding twice the multiplicand (left shift by one bit). Only the 5th partial product retains its original 8-bit width (because the high-order padding of two 0s means that, according to the encoding, the partial product is either set to 0 or is the multiplicand). Now, assuming all partial products are negative, as shown... Figure 2 As shown in (b), the high bits are all 1s at this point. Furthermore, according to the booth algorithm table, the data needs to be negative (inverted plus one). This can be achieved by pre-combining the partial product addition array: where x represents the two's complement plus one correction bit corresponding to subtracting twice the multiplicand (-2A), and y represents the two's complement plus one correction bit corresponding to subtracting the multiplicand (-A). If redundant consecutive sign extension bits are directly retained for partial product addition calculations, it will not only occupy additional circuit resources, but the frequent signal flipping generated by consecutive logic 1s during addition will also significantly increase circuit power consumption. Here, x0y0~x3y3 represent the first to fourth partial products, respectively.

[0065] The sign bit of the partially product in the case of all negative partial products is pre-processed, as shown in Figure 2(c). Comparing the two cases of all positive and all negative partial products and their corresponding schematic diagrams, it can be found that the difference lies in whether the lower bits are corrected by adding 1 and whether the higher bits are extended by an additional 2 bits. The difference in the lower bits is that variables x and y are introduced to distinguish the positive and negative partial products. When x and y are both set to 0, the corresponding partial product is in the all-positive state. At the same time, a variable S is added to the higher bits for characterization: when the overall partial product is negative, variable S is set to 1, and when the overall partial product is positive, variable S is set to 0, as shown in Figure 2(d), which is the simplified unsigned partial product addition circuit.

[0066] For signed multiplication, there are the following differences. First, there is the number of partial products. Because in Radix-4 Booth encoding, the sign bit is extended in the high-order bits of the signed number, the encoding of the 5th partial product can only be 000 or 111. In the Radix-4 Booth partial product selection table, these are both set to 0. Therefore, 8-bit signed number multiplication theoretically has only 4 partial products. However, because the 4th partial product may be negative, it will bring up the low-order bits. Therefore, the diagram still shows 5 partial products, but the high-order bits of the 5th partial product are all 0.

[0067] Furthermore, in unsigned number operations, the sign of the partial product is determined solely by the Booth encoding result; however, when switching to signed number operations, the data involved in the operation can carry a negative sign. Previously, in the analysis assuming all partial products were negative, negative partial products were extended with all 1s for the sign bit, and redundant sign extension bits were pre-simplified. Therefore, the original variable S was no longer suitable for the requirements of signed multiplication. To address this, a new variable E was added. Variable E takes the value of the XOR result of the multiplicand's sign and the multiplier's encoded sign. Its physical meaning is: it represents a negative attribute only when exactly one of them is negative; if both are negative, then negative negative equals positive, equivalent to an addition operation. A schematic diagram of the partial product summation in signed 8-bit multiplication is shown in Figure 2(e).

[0068] After completing Radix-4 Booth encoding, the partial product size of the 8-bit multiplier has been compressed from the traditional 8 rows to 5 rows, while introducing sign-extended S and E variables. Comparing signed and unsigned 8-bit multiplication partial product adder circuits, the difference lies in the S and E variables. In specific circuit implementations, the corresponding sign processing path can be switched via a signed / unsigned mode selection control signal, thus achieving unified support for both signed and unsigned modes within the same compression framework.

[0069] As another preferred embodiment of the present invention, this embodiment provides a detailed description of the layout architecture of the multi-precision parallel multiplication array structure and its mapping relationship with data precision. The row direction of the two-dimensional multiplication array structure is configured to support parallel computation of first-precision data, and the column direction is configured to support parallel computation of second-precision data, wherein the bit width of the first-precision data is lower than that of the second-precision data. Specifically, this embodiment breaks away from the traditional single-dimensional computation mode of arrays, innovatively utilizing the two orthogonal dimensions of the two-dimensional multiplication array structure to respectively carry computation tasks with different bit width precisions. Since the computation of low-bit-width data requires fewer basic multiplication units, and more independent computation groups can be divided in the row direction, configuring the row direction to support parallel computation of low-bit-width (first-precision) data can maximize parallel throughput. Conversely, the computation of high-bit-width data requires more cascaded basic multiplication units, and the column direction can provide longer cascaded paths; therefore, the column direction is configured to support parallel computation of high-bit-width (second-precision) data. This orthogonal row-column reuse architecture allows the same set of physical hardware resources to be dynamically reorganized in the row or column direction according to different data precision, which greatly improves the utilization efficiency of hardware resources and computing flexibility.

[0070] In one specific implementation, the two-dimensional array is 9 rows and 4 columns. The two-dimensional multiplication array structure is configured such that, in the column direction, multiple basic multiplication units are organized to support parallel computation of multiple sets of second-precision data; and in the row direction, multiple basic multiplication units are organized to support parallel computation of multiple sets of first-precision data. Combined with... Figure 1 and Figure 5 As shown, the 9x4 array size is not arbitrarily set, but rather an optimal solution derived through rigorous calculation based on the decomposition requirements of mainstream computational precision. The following details how this array size adapts to mainstream precision: For second-precision data, taking FP32 (single-precision floating-point number) as an example, the mantissa portion of FP32 has a bit width of 24 bits after recovering the hidden bits. Based on the default bit width of 8 bits for the basic multiplication unit, the multiplier decomposition and mapping module decomposes the 24-bit operand into three 8-bit data segments: high, medium, and low. , , According to the principle of multiplication expansion, multiplying two 24-bit numbers requires nine 8-bit sub-multiplication operations, that is:

[0071] Similarly:

[0072] The multiplication expansion is:

[0073]

[0074]

[0075] It is evident that this corresponds precisely to Figure 1 Each column of the array contains nine basic multiplication units. Therefore, the column direction of the array is configured to support 24-bit precision multiplication, with each column capable of independently performing the multiplication of a set of FP32 mantissas. Since the array has four columns, this structure can simultaneously support the parallel computation of four sets of FP32 data in the column direction, meeting the requirements of high-precision scientific computing.

[0076] For first-precision data, taking FP16 (half-precision floating-point number) as an example, the mantissa portion of FP16, after recovering the hidden bits, has a width of 11 bits. To adapt to the 8-bit basic computation unit, it is usually extended to 16 bits for processing. The multiplier decomposition and mapping module decomposes the 16-bit operand into two data segments: the high 8 bits and the low 8 bits; the high 8 bits... , 8 bits , According to the principle of multiplication expansion, multiplying two 16-bit numbers requires four 8-bit sub-multiplication operations, that is:

[0077] This corresponds precisely to the four basic multiplication units contained in a row of the array. Therefore, the row direction of the array is configured to support 16-bit precision multiplication operations, and each row can independently complete the multiplication task of a set of FP16 mantissas. Since the array has nine rows, this structure can simultaneously support the parallel computation of nine sets of FP16 data in the row direction, greatly improving the computational throughput in low-precision scenarios such as artificial intelligence inference.

[0078] Furthermore, for 8-bit integer data such as INT8, its bit width directly matches the preset bit width of the basic multiplication unit, and it can be directly mapped to any basic multiplication unit for calculation without decomposition. In this case, the 36 basic multiplication units in the array can work in full parallel, providing the highest computational parallelism.

[0079] This embodiment constructs an array architecture with extremely strong "defense depth" through the aforementioned row and column configuration logic. This design not only precisely matches the computational requirements of mainstream precisions such as FP32 and FP16, avoiding waste of hardware resources, but more importantly, it achieves dynamic scheduling of computational resources between tasks of different precisions through orthogonal reuse of row and column directions. For example, in artificial intelligence training scenarios, high-precision gradient calculation (FP32) can be performed simultaneously using the column direction, while low-precision activation value calculation (FP16) can be performed using the row direction, thus achieving efficient parallel processing of mixed precisions on a single array structure.

[0080] To facilitate understanding of the multiplier decomposition and mapping module provided in this embodiment, the following further explanation of the multiplier decomposition and mapping module is provided: In this embodiment, the multiplier decomposition and mapping module includes sub-modules such as 24-bit multiplier decomposition and mapping, 16-bit multiplier decomposition and mapping, and 8-bit multiplier mapping, which correspond to the 24-bit multiplication, 16-bit multiplication, and 8-bit multiplication completed by the designed parallel multiplication array result (reusable multiplier).

[0081] Furthermore, for the 24-bit multiplier decomposition and mapping submodule, this unit is used to process the multiplier decomposition and data allocation of 24-bit data (24-bit integers or the mantissa of FP32).

[0082] Taking the FP32 single-precision floating-point number under the IEEE 754 (Institute of Electrical and Electronics Engineers) specification as an example, the original 23-bit mantissa has a mantissa width of 24 bits after restoring the most hidden bit. This unit will decompose the two 24-bit input data A and B into three 8-bit data segments according to the high, middle, and low bits, for example, the high 8 bits. 8th 8 bits ,Right now:

[0083] In the formula, A is a 24-bit multiplicand. The high 8 bits of the 24-bit multiplicand The middle 8 bits of the 24-bit multiplicand The lower 8 bits of the 24-bit multiplicand.

[0084] Similarly:

[0085] In the formula, B is a 24-bit multiplier. The high 8 bits of the 24-bit multiplier The middle 8 bits of the 24-bit multiplier The lower 8 bits of the 24-bit multiplier.

[0086] The multiplication expansion is:

[0087]

[0088]

[0089] As can be seen from the multiplication expansion, a total of 9 8-bit sub-multipliers are needed to implement 24-bit multiplication, which corresponds exactly to a whole column of resources in this parallel multiplication array structure. In addition, each column has a compression and summation module at the end, which can quickly compress and sum the 9 partial products generated by the sub-multipliers to obtain the multiplication result of the original two 24-bit mantissas.

[0090] Explained, implementing 24-bit multiplication requires a series (9) of 8-bit sub-multipliers. Their mapping relationship simply needs to ensure that the multipliers and multiplicands in the 24-bit multiplication expansion are multiplied together and sent to the same sub-multiplier, such as... Figure 8 The data allocation mapping shown decomposes the data based on the resources of one column (9) of sub-multipliers in a 9×4 array multiplier. and Sent to the first sub-multiplier and Sent to the second sub-multiplier and The result is sent to the third sub-multiplier... and so on, ensuring that all element operations in the multiplication expansion are performed. Finally, the nine partial products of the sub-multipliers are quickly compressed and summed to obtain the final result.

[0091] A single FP32 mantissa multiplication requires 9 sub-multipliers, and since there are 4 columns of the same resources in the multiplication array, 4 sets of FP32 mantissa multiplications can be completed.

[0092] Furthermore, regarding the 16-bit multiplier decomposition and mapping module, this unit is used to perform multiplier decomposition and data allocation for multiplication of 16-bit data (the mantissa portion of INT16 or FP16). The half-precision floating-point number FP16 originally has a 10-bit mantissa width of 11 bits after restoring the highest hidden bit. However, a single 8-bit multiplier cannot process 11-bit data. To adapt to the 8-bit decomposition structure, it is extended to 16 bits, then divided into high 8 bits and low 8 bits. The numerical decomposition principle is the same as that of 24-bit multiplier decomposition.

[0093] For example, two input 16-bit data A and B, their high 8 bits... , 8 bits , The multiplication expansion is:

[0094] In the formula, A is a 16-bit multiplicand, B is a 16-bit multiplier, and... and These are the 16-bit multiplicand and the high 8 bits of the multiplier, respectively. and These are the lower 8 bits of the 16-bit multiplicand and multiplier, respectively. Multiplying by shifting the high and low bits requires four 8-bit sub-multipliers, which exactly corresponds to the multiplier resources in one row of the multiplication array. Combined with the compression and summation module at the end of each row, the multiplication result of the original two 16-bit data can be obtained.

[0095] Implementing 16-bit multiplication requires one row (four) of 8-bit sub-multipliers. Their mapping relationship should be such that the multipliers and multiplicands in the 16-bit multiplication expansion are sent to the same sub-multiplier. Figure 9 The alternative data allocation mapping shown decomposes the data based on the resources of one row (4) of sub-multipliers in a 9×4 array multiplier. and Sent to the first sub-multiplier and Sent to the second sub-multiplier and Sent to the third sub-multiplier and The result is sent to the fourth sub-multiplier to ensure that all element operations in the multiplication expansion are performed. Finally, the four partial products of the sub-multipliers are quickly summed to obtain the final result.

[0096] A single FP16 mantissa multiplication requires 4 sub-multipliers, and there are 9 rows of the same resources in the multiplication array, so theoretically a maximum of 9 sets of FP16 mantissa multiplications can be completed at once.

[0097] Furthermore, the 8-bit multiplier mapping unit only needs to distribute the data to the corresponding multiplier, such as the mantissa of an 8-bit integer (INT8) and a brain-precision floating-point number BF16 (the mantissa width of BF16 is 7 bits, which becomes 8 bits after restoring the highest hidden bit), which exactly matches the input width of a single 8-bit sub-multiplier unit, so there is no need for decomposition.

[0098] Because the 9×4 array multiplier has a total of 36 8-bit sub-multiplier resources, it can theoretically support a maximum of 36 INT8 or BF16 mantissa multiplications. However, due to the bandwidth limitation of data transmission, the computing resources can rarely be fully utilized.

[0099] For example, the multi-precision floating-point calculation unit originally supported in this embodiment has an input width of 128 bits for a single data item. This 128-bit width can accommodate 4 FP32, or 8 groups of FP16, or 8 groups of BF16, or 16 groups of INT8. After mantissa processing of the floating-point data, the data is packaged and sent to a 9×4 array multiplier for operation. The 9×4 array multiplier first needs to unpack the input data, and the unpacking method is as follows: Figure 10 As shown.

[0100] Due to the 128-bit transmission bandwidth, for FP32, since the input consists of 4 sets of FP32 data, and the 9×4 array multiplier's computational resources are just enough to complete the 24-bit mantissa multiplication of 4 sets of FP32, all the resources of the reusable array multiplier are utilized. For FP16, the input consists of 8 sets of FP16 data, and the 9×4 array multiplier's computational resources can complete the 11-bit mantissa multiplication of 9 sets of FP16, so one row of sub-multiplier resources will be unused in the data-to-sub-multiplier mapping. For BF16, the input consists of 8 sets of BF16 data, requiring 8 8-bit multipliers, and the entire array has 36 8-bit multipliers, so 8 sub-multipliers can be arbitrarily selected for mapping (e.g., two rows of sub-multipliers). For INT8, the input consists of 16 sets of 8-bit integer data, so 16 sub-multipliers can be arbitrarily selected for mapping (e.g., four rows of sub-multipliers).

[0101] Based on the above embodiments, this embodiment provides a detailed description of the specific structure and working principle of the compressed summation module. The compressed summation module includes a multi-level compression architecture, configured to compress and normalize the irregularly distributed operation results output by the basic multiplication unit into two lines of output data through multi-level compression.

[0102] Specifically, combined Figure 3 , Figure 6 and Figure 7 As shown, because the basic multiplication unit uses Radix-4 Booth encoding, although the number of partial products is reduced, the height of the generated partial product matrix in the column direction is not uniform, exhibiting an irregular distribution. For example, some weight bits correspond to a large number of bits, while others correspond to a small number of bits. If the traditional row-by-row addition method is used, the number of adder stages will increase linearly with the number of bits, resulting in excessively long critical path delays and severely limiting the operation speed of the multiplier. This embodiment introduces a multi-stage compression architecture, aiming to utilize the logic characteristics of the compressor to quickly "compress" these irregular multiple inputs into two regular rows of outputs (usually a row of sum bits and a row of carry bits). Finally, only one adder is needed to obtain the final result, thereby significantly reducing the logic depth.

[0103] As a specific implementation, the multi-stage compression architecture includes at least one of a 4-2 compressor, a 3-2 compressor, and a half-compressor. These three types of compressors each have their own characteristics in terms of circuit complexity and compression efficiency. This embodiment achieves efficient processing of irregular partial products by flexibly combining these three types of compressors.

[0104] Among them, the half-compressor, also known as a half-adder, compresses two inputs into one output and one carry; the 3-2 compressor is essentially equivalent to a full adder, compressing three equally weighted input bits into one output bit and one carry; the 4-2 compressor further extends this, compressing four equally weighted input bits and one low-order carry input into one sum bit and two carry outputs. The logic expressions of each compressor are as follows, and the schematic diagram is shown below. Figure 4 As shown in (a), (b), and (c); Figure 4 In this context, XOR represents the exclusive OR gate; AND represents the AND gate; and OR represents the OR gate.

[0105] Semi-compressor logic expression:

[0106]

[0107] 3-2 The logic expression of the compressor:

[0108]

[0109] 4-2 The logical expression of the compressor:

[0110]

[0111]

[0112] In the formula, , , , Each element is a single element to be summed (to be compressed) (i.e., the compressor input bit). For low-order carry input; sum is the summation result of the current compressor (sum bit output). and All are carry results from the current compressor ( (Appears only in 4-2 compressors), i.e., carry-out output; It is the XOR operator.

[0113] Because the optimized partial product addition circuit is irregular, a suitable compressor needs to be actively selected based on the number of data in each column during two-stage compression. For example, the first stage of compression prioritizes using a 4-2 compressor to efficiently compress columns with a large number of bits, reducing the number of operands while controlling the logic depth. For columns with 3 bits per column, a 3-2 compressor or a half-compressor is used for local merging. The second stage of compression further reduces the column height based on the first stage's compression result, so that the entire partial product matrix is ​​regularized into a two-row output format (Sum / Carry) after two-stage compression. Finally, the two rows of data are added to obtain the final result. The compression circuit is as follows: Figure 3 As shown.

[0114] The compression strategy in this embodiment is not a simple stacking, but rather a dynamic adaptation based on the distribution characteristics of the partial product matrix. Combined with... Figure 3 As shown, within the basic multiplication unit, the partial product matrix after Booth encoding may exhibit a shape that is tall in the middle and short at both ends. The compression and summation module prioritizes deploying 4-2 compressors in regions with higher column heights to quickly reduce the height; it uses 3-2 compressors in regions with moderate column heights; and it uses half-compressors in regions with lower column heights. Through this hierarchical and categorized compression strategy, the entire irregular partial product matrix is ​​quickly normalized into two rows of output.

[0115] This embodiment provides a 9×4 multiplication array structure. The shared multiplication array in this design consists of 36 8-bit sub-multipliers and several compression summation modules. The array spatial distribution is as follows: Figure 1As shown, with an 8-bit multiplier as the smallest sub-unit, there are 4 columns, each with 9 8-bit multipliers, totaling 36 sub-multipliers in 9 rows and 4 columns.

[0116] The principle of multiplier reconstruction is as follows: Figure 5 As shown, the input operands are divided into high-order and low-order bits according to their bit width. The partial product is calculated by multiple low-order multipliers, and the high-order bit multiplication result is achieved by shifting and adding.

[0117] Based on the multiplier reconstruction principle, the array structure can support the decomposition and multiplication of 4 groups of 24-bit data in the column-level structure, the decomposition and multiplication of 9 groups of 16-bit data in the row-level structure, and the multiplication of 36 groups of 8-bit data in the cell-level structure. Different precision modes are mapped in the array space to achieve physical resource sharing.

[0118] Combination Figure 6 and Figure 7 As shown, for high-bit-width multiplication (such as 16-bit or 24-bit) constructed using a two-dimensional multiplication array structure, the results output by multiple basic multiplication units need further aggregation. In this case, the compressed summation module also utilizes the aforementioned multi-level compression architecture to add the results of multiple sub-multiplications in a weighted, staggered manner. For example, in a 24-bit multiplication compressed summation structure, the results output by the nine basic multiplication units constitute a much larger partial product matrix. Through the layer-by-layer processing of the multi-level compression architecture, this matrix is ​​finally regularized into two rows of data before being fed into the final adder. This process effectively solves the bottleneck problem of result aggregation in multi-precision parallel computing, significantly reduces the number of addition stages and critical path latency, and ensures processing speed under high-precision computing conditions.

[0119] As another preferred embodiment of the present invention, this embodiment provides a detailed description of the specific workflow of the multiplier decomposition and mapping module. The multiplier decomposition and mapping module is configured to decompose operands into high-bit data segments and low-bit data segments according to their bit width, and then map the high-bit data segments and low-bit data segments to different basic multiplication units. Specifically, the multiplier decomposition and mapping module is the core scheduling unit for implementing multi-precision calculations. Its core idea is to "break down" high-bit-width operands into multiple data segments that adapt to the preset bit width of the basic multiplication units. It should be understood that although this embodiment mainly describes the decomposition into high-bit and low-bit segments, in practical applications, operands can be decomposed into two, three, or even more segments depending on the bit width difference of the target precision. The decomposed data segments carry information about the different weights of the original operands. When mapping to the basic multiplication units, their weight relationships must be strictly maintained to ensure the correctness of the subsequent summation result. This decomposition and mapping mechanism enables low-bit-width hardware resources to complete high-bit-width calculation tasks through combination and cooperation, greatly improving the hardware's versatility and reusability.

[0120] As a specific implementation, when the target precision is the first target precision, the multiplier decomposition and mapping module is configured to decompose the operands into high-order data segments, middle-order data segments, and low-order data segments, and map them to multiple basic multiplication units in the same column. Combined with... Figure 5 As shown, the mantissa multiplication with a first target precision of FP32 (single-precision floating-point number) is used as an example for explanation. The mantissa portion of FP32 has a width of 24 bits after recovering the hidden bits. Assuming the preset width of the basic multiplication unit is 8 bits, the multiplier decomposition and mapping module decomposes the 24-bit operand A into the high 8-bit data segment. 8-bit data segment and the lower 8 bits of data Similarly, operand B is also decomposed into , , Expanding according to the multiplication principle, the result contains 9 product terms. These 9 product terms correspond to different weights, for example... It needs to be shifted left by 32 bits. The lowest weights require no shifting. The multiplier decomposition and mapping module maps the computation of these 9 product terms to 9 basic multiplication units in the same column of a 2D array. Since the basic multiplication units in the same column are arranged vertically in physical position, their output naturally forms a vertical weight distribution, which facilitates staggered addition by the subsequent compression and summation module. This mapping method makes full use of the abundant resources in the column direction and perfectly adapts to the large number of basic units required for high-bit-width computations.

[0121] As another specific implementation, when the target precision is a second target precision, the multiplier decomposition and mapping module is configured to decompose the operand into high-order data segments and low-order data segments, and map them to multiple basic multiplication units in the same row. Taking mantissa multiplication with a second target precision of FP16 (half-precision floating-point number) as an example, the mantissa portion of FP16 has a width of 11 bits after recovering the hidden bits. To adapt to the 8-bit basic multiplication unit, it is usually extended to 16 bits for processing. The multiplier decomposition and mapping module decomposes the 16-bit operand into high-order 8-bit data segments. and the lower 8 bits of data Similarly, operand B is decomposed into and After expansion according to the multiplication expansion principle, there are a total of 4 product terms. The calculation of these 4 product terms is mapped to 4 basic multiplication units in the same row of the 2D array. Since there are 4 basic units in the same row, it perfectly satisfies the requirement of a set of FP16 mantissa multiplications. This mapping method takes advantage of the large number of basic units in the row direction, allowing the array to process multiple sets of low-bit-width data in parallel, greatly improving the computational throughput.

[0122] For example, the brain-precision floating-point number BF16 has a mantissa width of 7 bits, which becomes 8 bits after recovering the most significant hidden bit, perfectly matching the input width of a single 8-bit sub-multiplication unit. Another example is quantized data INT8, which is itself 8 bits wide and can be directly mapped to a single 8-bit sub-multiplication unit. Furthermore, quantization is divided into symmetric and asymmetric quantization, meaning that after quantization, it becomes both signed and unsigned INT8, both of which are supported by the sub-multiplication unit.

[0123] In summary, the designed reusable multiplier array contains 36 8-bit sub-multipliers and several compression summation modules. A single calculation can complete up to 4 sets of 24-bit multiplications, or 9 sets of 16-bit multiplications, or 36 sets of 8-bit multiplications.

[0124] This embodiment achieves deep coupling between the algorithm and the hardware architecture through the aforementioned refined decomposition and mapping logic. The multiplier decomposition and mapping module not only completes the physical partitioning of the data, but more importantly, it completes the spatial mapping of logical weights. For computational tasks of different precisions, the module can automatically identify their bit width characteristics, select the optimal decomposition strategy (three-segment or two-segment) and mapping path (column direction or row direction), thereby achieving multi-precision adaptive and efficient computation on a unified hardware array.

[0125] For example, this embodiment also provides a multi-precision parallel multiplication operation method. This method is based on the multi-precision parallel multiplication array structure in the above embodiment, and includes the following steps: Step S100: Decompose the operands of the target precision into multiple data segments, and map the multiple data segments to the corresponding basic multiplication units in the two-dimensional array.

[0126] Specifically, this step is the data preprocessing and resource scheduling stage before computation. After receiving the operands with the target precision, it is first determined whether the bit width of the operands is greater than the preset bit width of the basic multiplication unit. If it is greater, the high-bit-width operands are divided into several data segments matching the preset bit width according to the preset decomposition rules. For example, if the target precision is 24 bits and the preset bit width is 8 bits, the operands are decomposed into three 8-bit data segments: high, medium, and low. Subsequently, based on the weight position of the data segments in the original operands, they are mapped to specific basic multiplication units in the two-dimensional array. This mapping is not random, but based on the expansion principle of multiplication operations to ensure that each data segment can be allocated to the correct computation position, preparing for subsequent parallel computation. It should be understood that the number of segments decomposed and the mapping positions depend on the ratio between the target precision and the preset bit width, and this embodiment does not specifically limit it.

[0127] Step S200: Perform a multiplication operation of a preset bit width on the data segment through the basic multiplication unit to obtain multiple operation results.

[0128] Specifically, this step is the core stage of parallel computing. Data segments mapped to each basic multiplication unit in the two-dimensional array serve as the input operands for that unit. Each basic multiplication unit independently and in parallel performs multiplication operations of a preset bit width (e.g., 8 bits). Since the two-dimensional array contains multiple basic multiplication units, multiplication operations on multiple data segments can be performed simultaneously, greatly reducing the overall computation time. The output result of each basic multiplication unit is essentially a partial product or intermediate result in the target precision multiplication operation. Although these results have been calculated, they cannot be directly merged into the final result because they correspond to different weight bits and need to be processed in the next step.

[0129] Step S300: Compress and sum the multiple calculation results to obtain the multiplication result with the target precision.

[0130] Specifically, this step is the aggregation and regularization stage of the results. Since the multiple operation results output in step S200 have misaligned weights and are numerous, using the traditional row-by-row addition method would lead to excessive latency. This step employs a compressed summation method, utilizing a compression tree structure composed of compressors (such as 4-2 compressors, 3-2 compressors, etc.) to quickly compress multiple operation results into two rows of output (such as the sum and carry bits), which are then added together by an adder to obtain the final multiplication result. This process effectively reduces the number of addition stages and lowers the critical path latency. Through the above steps S100 to S300, this embodiment achieves a complete closed loop from data decomposition and parallel computation to result aggregation, enabling low-bit-width computing resources to efficiently and collaboratively complete high-bit-width multiplication tasks, significantly improving computational efficiency and resource utilization while ensuring computational accuracy.

[0131] For example, this embodiment applies the above-described multi-precision parallel multiplication array structure to a specific artificial intelligence accelerator computing scenario to further illustrate its technical effects and commercial value in practical applications.

[0132] In artificial intelligence neural network inference scenarios, computational tasks often involve a large number of low-precision integer operations, such as multiplication of INT8 quantized data. Traditional GPUs or accelerators are often configured with fixed-width multipliers, resulting in a significant waste of resources when processing INT8 data. However, in the array structure of this embodiment, since the basic multiplication unit is configured with a preset width of 8 bits and supports switching between signed and unsigned operation modes, the array can handle such tasks with the highest resource utilization.

[0133] Specifically, when the AI ​​accelerator receives a neural network inference instruction, the multiplier decomposition and mapping module identifies the input operand as INT8 type. At this point, the bit width of the operand perfectly matches the preset bit width of the basic multiplication unit, eliminating the need for data decomposition. The two-dimensional multiplication array structure configures all basic multiplication units (e.g., 36 units in 9 rows and 4 columns) in the two-dimensional array as independent computing resources. Each basic multiplication unit responds to a configuration signal, switching to the corresponding signed or unsigned operation mode of INT8. Thus, within a single clock cycle, the array can complete 36 sets of INT8 multiplication operations in parallel. Compared to the traditional solution that uses 32-bit multipliers for time-division multiplexing of INT8 data, the solution provided in this embodiment increases the inference throughput by several times within the same chip area, greatly satisfying the high-concurrency, low-latency inference requirements of edge AI devices.

[0134] In another common AI training or high-precision inference scenario, the processing of FP16 (half-precision floating-point) data is often involved. The mantissa portion of FP16, after recovering the hidden bits, has a width of 11 bits, which is typically expanded to 16 bits for calculation. This is where the multiplier factorization and mapping module comes into play.

[0135] Specifically, the multiplier decomposition and mapping module decomposes the 16-bit operand into high-order and low-order data segments. The two-dimensional multiplication array structure, based on the architecture of the above embodiment, configures the row direction of the two-dimensional array to support FP16 parallel computation. Since each row contains four basic multiplication units, it precisely meets the requirement of four 8-bit sub-multiplication operations required for a set of FP16 mantissa multiplications. Therefore, each row in the array is assigned to process a set of FP16 data. For a 9x4 array, nine sets of FP16 multiplications can be processed in parallel at a time. This row-direction parallel mapping mechanism fully utilizes the array's spatial resources, enabling efficient pipelined execution of gradient and activation value calculations during AI training. Simultaneously, because the basic multiplication units employ Radix-4 Booth encoding and S / E variable optimization, power consumption and latency are effectively controlled while ensuring FP16 computation accuracy, making it particularly suitable for data center AI accelerator cards with extremely high energy efficiency requirements.

[0136] In scientific computing or high-precision neural network training scenarios, FP32 (single-precision floating-point number) is the mainstream data format. The mantissa portion of FP32 has a bit width of 24 bits, which places high demands on the bit width of the multiplier. In this embodiment, the two-dimensional multiplication array structure configures the column direction of the two-dimensional array to support parallel FP32 computation.

[0137] Specifically, the multiplier decomposition and mapping module decomposes the 24-bit mantissa operand into high-order, middle-order, and low-order data segments, and maps them to multiple basic multiplication units in the same column. According to the multiplication expansion principle, a set of FP32 mantissa multiplications requires nine 8-bit sub-multiplication operations, which corresponds precisely to the nine basic multiplication units in one column of the array. Therefore, each column of the array is assigned to process a set of FP32 data. For a 9x4 array, four sets of FP32 multiplications can be processed in parallel at a time. This column-direction cascaded mapping breaks the limitation that low-order width units cannot process high-order width data, allowing the same hardware resources to smoothly support high-precision scientific computing without refactoring the circuit structure. The compression and summation module then performs multi-level compression on the multiple intermediate results output in the column direction to quickly obtain the final FP32 multiplication result, ensuring the data path efficiency of high-precision computing.

[0138] In summary, this embodiment demonstrates the practical performance of a multi-precision parallel multiplication array structure in an artificial intelligence accelerator. Whether facing low-concurrency INT8 inference tasks, high-throughput FP16 training tasks, or high-precision FP32 scientific computing tasks, this array structure can achieve dynamic adaptation and efficient reuse of hardware resources through the coordinated configuration of data decomposition mapping and array arrangement. Compared to existing technologies that design independent multipliers for different precisions, this invention significantly reduces chip area overhead and demonstrates excellent throughput and flexibility in multi-precision mixed computing scenarios. In summary, this invention provides a multi-precision parallel multiplication array structure, which has the following advantages compared to existing multipliers: First, it significantly improves resource utilization: due to the use of 8-bit sub-unit multiplexing, the same hardware can support multiplication implementations of various precisions, avoiding the need to repeatedly design multipliers with different bit widths.

[0139] Second, it improves reconfigurability and flexibility: through fine-grained (8-bit) partitioning, it supports multiple precision combinations and can be adapted to different application scenarios.

[0140] Third, it has good scalability: the two-dimensional array structure allows the multiplication implementation to be extended to higher bit widths without changing the basic unit structure.

[0141] Fourth, strong parallel computing capability: multiple sub-operation units perform parallel computing, improving the overall throughput.

[0142] Fifth, the structure is regular and easy to implement: the array structure is conducive to layout design and hardware implementation.

[0143] The above embodiments are merely one of the implementation methods for achieving the technical solution of the present invention. The scope of protection claimed by the present invention is not limited to this embodiment, but also includes any variations, substitutions and other implementation methods that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention.

Claims

1. A multi-precision parallel multiplication array system, characterized in that, include: Two-dimensional multiplication array structure; The two-dimensional multiplication array structure includes multiple basic multiplication units with identical structures; Multiple basic multiplication units are arranged in a two-dimensional array and constructed into multipliers of different bit widths by combining row-level and column-level arrays; each basic multiplication unit serves as the minimum multiplexing unit and is used to perform multiplication operations of a preset bit width. The multiplier decomposition and mapping module is used to receive operands of the target precision, decompose the operands into multiple data segments according to the bit width of the operands, and map the multiple data segments to the corresponding basic multiplication units in the two-dimensional multiplication array structure. The compression and summation module is used to perform hierarchical parallel compression and summation on the operation results output by the basic multiplication unit to obtain the multiplication result with the target precision. The row direction of the two-dimensional multiplication array structure is configured to support parallel computation of first-precision data, and the column direction of the two-dimensional array is configured to support parallel computation of second-precision data; wherein the bit width of the first-precision data is lower than the bit width of the second-precision data. The compression and summation module includes a multi-level compression architecture, configured to compress and normalize the irregularly distributed operation results output by the basic multiplication unit into two rows of output data through hierarchical parallel compression, and add the two rows of output data to obtain the multiplication result with the target precision; the compression and summation module dynamically selects the compressor type according to the actual height of each column; The multi-stage compression architecture includes one or more of a 4-2 compressor, a 3-2 compressor, and a half-compressor; wherein: The 4-2 compressor is configured to compress four equally weighted input bits and one low-carry input into one sum bit and two carry outputs; The 3-2 compressor is configured to compress three equally weighted input bits into one output bit and one carry; The semi-compressor is configured to compress two inputs into one output and one carry.

2. The multi-precision parallel multiplication array system according to claim 1, characterized in that, The basic multiplication unit is also used to switch between signed number multiplication operation mode and unsigned number multiplication operation mode based on the received configuration signal.

3. A multi-precision parallel multiplication array system according to claim 2, characterized in that, The basic multiplication unit includes: The encoding subunit is used to encode operands of a preset bit width using the radix-4 Booth encoding algorithm to generate a partial product; The symbol processing subunit specifically includes: a first variable generation circuit for generating a first variable for processing the sign bit extension; and a second variable generation circuit for generating a second variable for uniformly processing the partial product sign in both signed and unsigned multiplication operation modes.

4. A multi-precision parallel multiplication array system according to claim 1, characterized in that, The basic multiplication unit adopts an 8-bit sub-multiplier or a 4-bit sub-multiplier; the two-dimensional multiplication array structure adopts a regular array of 9 rows and 4 columns, integrating a total of 36 basic multiplication units.

5. A multi-precision parallel multiplication array system according to claim 1, characterized in that, The execution logic expression of the 4-2 compressor is as follows: The execution logic expression of the 3-2 compressor is as follows: The execution logic expression of the semi-compressor is as follows: In the formula, , , , These are the compressor input bits; For carry-in from the least significant bit; For compressor and bit output; Outputting the carry from the middle; For carry-out output; It is the XOR operator.

6. A multi-precision parallel multiplication array system according to claim 1, characterized in that, When the target precision is the first target precision, the multiplier decomposition and mapping module is configured to decompose the operand into high-bit data segments, middle-bit data segments and low-bit data segments according to the bit width of the basic multiplication unit, and map them to multiple basic multiplication units in the same column. When the target precision is the second target precision, the multiplier decomposition and mapping module is configured to decompose the operand into high-bit data segments and low-bit data segments according to the bit width of the basic multiplication unit, and map them to multiple basic multiplication units in the same row. The bit width of the first target precision is higher than that of the second target precision.

7. A multi-precision parallel multiplication method, based on the multi-precision parallel multiplication array system according to any one of claims 1-6, characterized in that, include: The multiplier decomposition and mapping module receives operands of the target precision, decomposes the operands into multiple data segments according to the bit width of the operands, and maps the multiple data segments to the corresponding basic multiplication units in the two-dimensional multiplication array structure. The basic multiplication unit, employing a two-dimensional multiplication array structure, performs multiplication operations of a preset bit width on multiple data segments and outputs the results. A compression and summation module is used to perform hierarchical parallel compression and summation on the operation results output by the basic multiplication unit to obtain the multiplication result with the target precision.

Citation Information

Patent Citations

  • Multi-granularity parallel storage system

    CN102541749A

  • Structured mixed bit-width multiplying method and structured mixed bit-width multiplying device

    CN102591615A