BP16 matrix multiplication and addition operation circuit based on risc-v architecture
Patent Information
- Application Number
- CN202610509196.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-17
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-04-17
AI Technical Summary
[0005]针对现有技术的不足,本申请提供了一种基于RISC-V架构的BP16矩阵乘加运算电路,可以解决BP16矩阵乘加运算效率低、硬件复杂度高以及特殊值处理不统一的问题
[0018]本申请提供的基于RISC-V架构的BP16矩阵乘加运算电路,包括输入模块、矩阵乘加运算模块、累加运算模块及输出模块,输入模块用于输入BP16格式的第一操作数、BP16格式的第二操作数及预设格式的第三操作数;矩阵乘加运算模块周期性地将第一操作数中预设行数的多个第一浮点元素与第二操作数中的第二浮点元素以矩阵形式进行行列乘法运算,生成第四操作数;累加运算模块将第四操作数与第三操作数进行累加计算,生成第五操作数;输出模块基于第五操作数输出目标操作数,解决了BP16矩阵乘加运算效率低、硬件复杂度高以及特殊值处理不统一的问题,显著提升了RISC-V平台在AI加速场景下的算力密度、能效比与数值鲁棒性。
Smart Images

Figure CN122064314B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of processor microarchitecture and artificial intelligence computing acceleration technology, and in particular to a BP16 matrix multiplication and addition circuit based on the RISC-V architecture. Background Technology
[0002] Matrix multiplication and addition are core operators for accelerating computation, scientific computing, and vector processing in artificial intelligence. Their performance directly limits the throughput and energy efficiency of deep learning model training and inference. With the rapid deployment of the RISC-V architecture in AI edge computing and customized accelerators, there is an urgent need for lightweight, high-density, and scalable floating-point computing hardware support.
[0003] BP16 (Brain Precision 16), as a 16-bit floating-point format, inherits the 8-bit exponent width of FP32. While maintaining a wide dynamic range, it significantly reduces data path bandwidth and storage overhead, and has become a key data type for RISC-VAI extended instruction sets (such as V extensions and Zfhmin).
[0004] However, current mainstream processors and accelerators still rely on general-purpose floating-point units or software simulation to implement BP16 matrix multiplication and addition, which has problems such as low computing power density, insufficient memory access bandwidth utilization, high latency, and poor integration of RISC-V native instructions. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this application provides a BP16 matrix multiplication and addition circuit based on the RISC-V architecture, which can solve the problems of low efficiency, high hardware complexity, and inconsistent handling of special values in BP16 matrix multiplication and addition operations.
[0006] In a first aspect, this application provides a BP16 matrix multiplication and addition operation circuit based on a RISC-V architecture, comprising: The input module is configured to input operands, which include a first operand, a second operand, and a third operand. The first operand includes multiple first floating-point elements in BP16 format, the second operand includes multiple second floating-point elements in BP16 format, and the third operand includes multiple third floating-point elements in a preset format. The matrix multiplication and addition module is connected to the input module and is configured to periodically perform row and column multiplication operations on multiple first floating-point elements in a preset row number in the first operand and second floating-point elements in the second operand in the form of a matrix to obtain a fourth operand, which includes multiple fourth floating-point elements in a preset format. The accumulation operation module connects the matrix multiplication and addition operation module and the input module, and is configured to accumulate multiple fourth floating-point elements with third floating-point elements to obtain a fifth operand, which includes multiple fifth floating-point elements in a preset format. The output module is connected to the accumulation module and is configured to output the target operand based on multiple fifth floating-point elements.
[0007] In one embodiment, the BP16 matrix multiply-accumulate circuit based on the RISC-V architecture further includes a preprocessing module; The input end of the preprocessing module is connected to the input module, and the output end of the preprocessing module is connected to the matrix multiplication and addition operation module. The preprocessing module is configured to preprocess the first operand, the second operand, and the third operand respectively to obtain the preprocessed first operand, the second operand, and the third operand.
[0008] In one embodiment, the exponents of the floating-point elements in the third and fourth floating-point elements after preprocessing are represented using signed two's complement.
[0009] In one embodiment, when the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element reach the minimum negative value represented by the two's complement, independent protection processing is performed on the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element.
[0010] In one embodiment, the preprocessing module includes a splitting module, a special data detection module, and a hidden bit recovery module; The splitting module is connected to the input module, the hidden bit recovery module, the matrix multiplication and addition module, and the accumulation module, respectively. The special data detection module is connected to the input module, the hidden bit recovery module, and the accumulation module, respectively. The hidden bit recovery module is connected to the matrix multiplication and addition module and the accumulation module, respectively. The splitting module is configured to split the first floating-point element, the second floating-point element, and the third floating-point element into bits respectively, so as to obtain the sign bit, exponent, and mantissa in the first floating-point element, the second floating-point element, and the third floating-point element. The special data detection module is configured to perform normalization detection on the mantissas of the first floating-point element, the second floating-point element, and the third floating-point element respectively, and obtain the normalization detection results. The hidden bit recovery module is configured to recover the hidden bit of the mantissa in the first, second, and third floating-point elements based on the normalization detection results.
[0011] In one embodiment, the preprocessing module further includes a leading zero detection module, a mantissa normalization module, and an exponent adjustment module; Among them, the leading zero detection module is connected to the matrix multiplication and addition operation module, the accumulation operation module, the hidden bit recovery module and the mantissa normalization module respectively; the mantissa normalization module is connected to the matrix multiplication and addition operation module and the accumulation operation module respectively; and the exponent adjustment module is connected to the splitting module, the leading zero detection module and the matrix multiplication and addition operation module respectively. The leading zero detection module is configured to perform leading zero detection on the denormalized mantissas in the first floating-point element, the second floating-point element, and the third floating-point element respectively, and obtain the first leading zero detection result; The mantissa normalization module is configured to normalize the unnormalized mantissas in the first floating-point element and the second floating-point element based on the first leading zero detection result. Both the exponent adjustment module and the mantissa normalization module are configured to process the exponent and the denormalized mantissa in the third floating-point element based on the first leading zero detection result, so as to obtain the preprocessed third operand.
[0012] In one embodiment, the accumulation module includes an exponent comparison module, an alignment shift module, and an addition module; The exponent comparison module is connected to the preprocessing module, the alignment and shifting module, and the output module, respectively; the addition module is connected to the preprocessing module, the alignment and shifting module, and the output module, respectively. The exponent comparison module is configured to filter the largest exponent from the exponents of the floating-point elements in the preprocessed third operand and the exponents of multiple fourth floating-point elements to obtain the exponent of the fifth floating-point element. The alignment shift module is configured to right-shift and align the mantissas of multiple fourth floating-point elements based on the maximum exponent, resulting in multiple right-shifted and aligned mantissas. The addition module is configured to add the mantissa of the floating-point element in the preprocessed third operand to multiple right-shifted mantissas to obtain the mantissa of the fifth floating-point element.
[0013] In one embodiment, the output module is configured to perform leading zero detection on the mantissas of a plurality of fifth floating-point elements to obtain a second leading zero detection result, and to normalize the mantissas of the plurality of fifth floating-point elements based on the second leading zero detection result to generate the mantissa in the target operand; to adjust the exponents of the plurality of fifth floating-point elements based on the second leading zero detection result to generate the exponent in the target operand; and to output the target operand.
[0014] In one embodiment, the matrix multiply-add operation module includes a multiplexer and a multiplication operation module; The multiplexer is connected to the input module and the multiplication module, respectively, and the multiplication module is connected to the accumulation module. The multiplexer is configured to periodically output multiple first floating-point elements of a preset number of rows in the first operand; The multiplication module is configured to perform matrix multiplication on multiple first floating-point elements in a preset row number in the first operand and second floating-point elements in the second operand to obtain multiple fourth floating-point elements.
[0015] In one embodiment, the matrix multiplication and addition module further includes a processing module; The processing modules are connected to the accumulation operation modules respectively; The multiplication module is configured to perform row and column multiplication on multiple first floating-point elements in a preset row number in the first operand and second floating-point elements in the second operand in the form of a matrix, to obtain multiple sixth floating-point elements; The processing module is configured to expand the mantissas of multiple sixth-point elements and align the exponents of multiple sixth-point elements to obtain a fourth operand.
[0016] Secondly, this application also provides a processor, including the BP16 matrix multiplication and addition circuit based on the RISC-V architecture provided in the first aspect.
[0017] Thirdly, this application also provides an electronic device, including the BP16 matrix multiply-accumulate circuit based on the RISC-V architecture provided in the first aspect or the processor provided in the second aspect.
[0018] The BP16 matrix multiplication-addition circuit based on the RISC-V architecture provided in this application includes an input module, a matrix multiplication-addition module, an accumulation module, and an output module. The input module is used to input a first operand in BP16 format, a second operand in BP16 format, and a third operand in a preset format. The matrix multiplication-addition module periodically performs matrix multiplication on multiple first floating-point elements in a preset number of rows of the first operand and the second floating-point elements in the second operand to generate a fourth operand. The accumulation module accumulates the fourth operand and the third operand to generate a fifth operand. The output module outputs the target operand based on the fifth operand. This solves the problems of low efficiency, high hardware complexity, and inconsistent handling of special values in BP16 matrix multiplication-addition operations, and significantly improves the computing power density, energy efficiency ratio, and numerical robustness of the RISC-V platform in AI acceleration scenarios. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1A first schematic block diagram of a BP16 matrix multiply-add operation circuit based on a RISC-V architecture provided for embodiments of this application; Figure 2 A second schematic block diagram of a BP16 matrix multiply-add operation circuit based on a RISC-V architecture provided for embodiments of this application; Figure 3 A first architecture diagram of the BP16 matrix multiplication and addition circuit based on RISC-V architecture provided for embodiments of this application; Figure 4 The second architecture diagram of the BP16 matrix multiplication and addition operation circuit based on the RISC-V architecture provided in the embodiments of this application is shown. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0023] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0024] It should also be further understood that the term “and / or” as used in this application specification and the appended claims is to mean any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0025] Furthermore, in this application, unless otherwise explicitly specified or limited in the embodiments, the terms "installation," "connection," "joining," and "fixing" appearing in the embodiments should be interpreted broadly. For example, a connection can be a fixed connection, a detachable connection, or an integral part; it can also be a mechanical connection, an electrical connection, etc. Of course, it can also be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication between two components, or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific implementation.
[0026] In related technologies, deep learning models are continuously increasing their demands for energy efficiency and throughput in low-bit floating-point operations, particularly in scenarios such as artificial intelligence inference and training, and high-performance vector computing. Currently, BF16 is often used for matrix multiplication and addition operations. Although it has high numerical accuracy, it has significant bottlenecks in terms of computing power density, power consumption control, and storage bandwidth. In particular, under the RISC-V architecture, the general-purpose floating-point pipeline is difficult to efficiently support large-scale parallel matrix operations in the BP16 format. Furthermore, the handling of denormalized numbers (Denorm), NaN, Inf, and exponential boundary anomalies is scattered across multiple logic units, resulting in high hardware complexity, high latency, and low energy efficiency.
[0027] Furthermore, during the accumulation process from BP16×BP16 to FP32, due to the large differences in the dynamic range of the exponent, the narrowness of the mantissa, and the tight coupling between rounding and normalization, precision loss or overflow errors are prone to occur, and there is a lack of a unified and predictable hardware co-processing mechanism.
[0028] To address this, this application provides a BP16 matrix multiplication and addition circuit based on the RISC-V architecture, which achieves efficient hardware implementation of BP16 matrix multiplication and addition operations, significantly improving computing power density and energy efficiency in AI scenarios. It natively supports the BP16 format, reducing data migration and bandwidth pressure. It is highly integrated with the RISC-V vector architecture, and performs unified and predictable hardware processing of BP16 special values and exponential boundaries. It is suitable for CPUs, AI accelerators, and heterogeneous computing platforms.
[0029] In the BP16 matrix multiplication and addition circuit based on the RISC-V architecture, three operands are received through the input module, including the first operand in BP16 format, the second operand in BP16 format, and the third operand in a preset format, thus realizing efficient processing of low-bit floating-point data.
[0030] This application utilizes a matrix multiplication and addition module to periodically perform matrix multiplication operations on multiple first floating-point elements in a preset number of rows of the first operand and second floating-point elements in the second operand, generating intermediate results. This solves the problem of the lack of dedicated hardware support for BP16 matrix multiplication and addition, significantly improving computational efficiency. The accumulation module adds the intermediate results to the third operand, achieving precision conversion from BP16×BP16 to FP32, effectively avoiding the accumulation of numerical errors due to insufficient precision. The output module outputs the target operand based on the accumulated results, ensuring the completeness and usability of the calculation results.
[0031] Furthermore, this application also uses a preprocessing module to perform format adaptation and pre-normalization on the input operands, solving the problem of handling special values such as denormalized numbers, NaN, and Inf in BP16 data, and ensuring the numerical consistency of the calculations. At the same time, the exponents of the floating-point elements in the third and fourth floating-point elements of the preprocessed operands are represented using signed two's complement, simplifying the exponent comparison logic, supporting wide dynamic range mapping from BP16 to FP32, and ensuring correct alignment even in the smallest exponent scenario.
[0032] Meanwhile, independent protection is performed on the minimum negative exponent, eliminating the risk of numerical collapse under the worst boundary and ensuring the determinism and verifiability of the calculation results in the entire input domain.
[0033] The collaborative work of the splitting module, special data detection module, and hidden bit recovery module enables unified parsing of the BP16 format, explicitly exposing hidden bits and avoiding loss of mantissa precision. The cooperation of the leading zero detection module, mantissa normalization module, and exponent adjustment module eliminates computational redundancy caused by denormalized numbers, unifies the exponent reference system, and ensures accurate accumulation alignment.
[0034] The combination of the exponent comparison module, alignment shift module, and addition operation module ensures that the numerical accuracy of the accumulation process is not degraded, supports multi-channel parallel accumulation, and is compatible with the GRS bit generation required by the RISC-V rounding mode.
[0035] Meanwhile, the output module can also perform leading zero detection, normalization shift, and dynamic exponent adjustment to ensure that the output strictly conforms to the IEEE 754 FP32 format, supports five RISC-V rounding modes, and realizes automatic setting of the exception flag.
[0036] The multiplexer and multiplication module work together to achieve high-throughput matrix multiplication, supporting periodic issuance of RISC-V vector instructions and reducing the temporary storage area for multiplication results. The processing module performs mantissa expansion and exponent alignment on the floating-point elements output by the multiplication module, eliminating the data format gap between the multiplication and accumulation modules and avoiding truncation errors caused by bit width mismatch.
[0037] This application achieves efficient hardware implementation of BP16 matrix multiplication and addition operations through the organic combination of the above-mentioned technical features, significantly improving the computing power density and energy efficiency ratio in AI scenarios. It natively supports the BP16 format, reduces data migration and bandwidth pressure, and is highly integrated with the RISC-V vector architecture. It provides unified and predictable hardware processing for BP16 special values and exponential boundaries, and is suitable for CPUs, AI accelerators and heterogeneous computing platforms.
[0038] Please see Figure 1 , Figure 1 This is a first schematic block diagram of a BP16 matrix multiplication and addition operation circuit based on the RISC-V architecture provided for embodiments of this application. Figure 1 As shown, this application provides a BP16 matrix multiplication and addition operation circuit based on RISC-V architecture, including: The input module 110 is configured to input operands, which include a first operand, a second operand, and a third operand. The first operand includes multiple first floating-point elements in BP16 format, the second operand includes multiple second floating-point elements in BP16 format, and the third operand includes multiple third floating-point elements in a preset format. The matrix multiplication and addition module 120 is connected to the input module 110 and is configured to periodically perform row and column multiplication operations on multiple first floating-point elements in a preset row number of the first operand and second floating-point elements in the second operand in the form of a matrix to obtain a fourth operand, the fourth operand including multiple fourth floating-point elements in a preset format. The accumulation operation module 130 is connected to the matrix multiplication and addition operation module 120 and the input module 110, and is configured to accumulate multiple fourth floating-point elements with third floating-point elements to obtain a fifth operand, the fifth operand including multiple fifth floating-point elements in a preset format; Output module 140 is connected to accumulation module 130 and is configured to output target operands based on multiple fifth floating-point elements.
[0039] Specifically, this application provides a hardware native adaptation approach for the RISC-V Vector Extension (RVV) instruction set. Specifically, it can take FP32+=BP16×BP16 as the basic computation paradigm, modularly decouple and periodically pipeline the input data stream, matrix operation structure, accumulation alignment logic and output regularization process, and achieve high throughput, low power consumption and strong robust BP16 matrix multiplication and addition hardware acceleration while ensuring IEEE 754 compatibility and RISC-V rounding semantic consistency.
[0040] The input module 110 can be understood as an interface circuit for receiving and buffering three sets of operand data paths from the RISC-V Vector Register File (VRF) or memory controller. Its input port width is adapted to a 512-bit bus (e.g., supporting 16 BP16 elements or 4 FP32 elements for parallel input). Internally, it contains a multi-level register set and timing alignment logic to meet the data tick requirements of vfmacc.vv type instructions in the RISC-VV extended instructions.
[0041] The input module 110 can provide a synchronous, stable, and bandwidth-matched data source for subsequent operations. It achieves back pressure control with the matrix multiplication and addition module 120 and the accumulation module 130 through handshake signals (valid / ready) to avoid data loss. At the same time, the input module 110 does not perform any format conversion or numerical processing, but only undertakes the functions of data transfer and clock buffering. Its structure can be set according to the actual chip process node and frequency target to determine the number and depth of registers. For example, it can be configured as a 2-level pipelined register or a 3-level cross-clock domain synchronization register.
[0042] Both the first and second operands are in BP16 format, which can be understood as... Figure 4 The opa and opb in the text.
[0043] The matrix multiplication and addition module 120 is a dedicated hardware unit for performing BP16×BP16 matrix multiplication. Its input terminal is connected to the first and second operands of BP16 output by the input module 110, and its output terminal generates the fourth operand.
[0044] The matrix multiplication and addition module 120 is configured to periodically perform row and column multiplication operations on multiple first floating-point elements in a preset number of rows in the first operand and second floating-point elements in the second operand in matrix form. The preset number of rows is 1, 2, or 4, corresponding to different matrix partitioning strategies. For example, in the (8×4)×(4×8) calculation scenario, 2 rows × 4 columns of BP16 input and 1 column × 4 rows of BP16 input are processed each cycle to generate 2 rows × 8 columns of FP32 intermediate results.
[0045] The matrix multiplication and addition module 120 includes a multiplexer and a multiplication submodule. Its function is to take the BP16 input and generate an extended intermediate result in the format Sgn(1)+Exp(10)+Man(24) after XORing the sign bit, adding and subtracting the exponents and compensating for the leading zeros, multiplying the mantissa by Booth-4 and restoring the hidden bits, so as to output the fourth operand.
[0046] Simultaneously, the matrix multiplication and addition module 120 and the input module 110 are coordinated and scheduled through address indexing and enable signals, starting a fixed number of pipeline operations after RISC-V vector instruction decoding; the connection relationship between the sub-modules in this module is as follows: the multiplexer outputs the selected BP16 row data to the multiplication sub-module, the multiplication sub-module outputs multiple fourth floating-point elements to the accumulation module 130, and thus the matrix multiplication and addition module 120 can complete 4 BP16 multiplications and 3 intermediate product additions in a single cycle, forming an FP32 precision accumulation intermediate term; its working result is a set of extended exponent range ([ [139,381]) and the fourth floating-point element with 24-bit mantissa precision. This result is not aligned with the third operand exponent, nor is normalization or rounding performed. It only provides a high-precision intermediate representation for subsequent accumulation.
[0047] The accumulation operation module 130 is a dedicated floating-point addition unit for performing FP32 precision accumulation and fusion. Its first input terminal is connected to the fourth operand output by the matrix multiplication and addition operation module 120, its second input terminal is connected to the third operand output by the input module 110, and its output terminal generates a fifth operand. This module is configured to accumulate multiple fourth floating-point elements with third floating-point elements, where the third operand is the initial accumulated value in FP32 format (i.e., the old accumulated sum), and the fourth operand is the extended FP32 intermediate result generated by BP16×BP16.
[0048] The accumulation operation module 130 may include an exponent comparison submodule, an alignment shift submodule, and an addition operation submodule. Its function is to realize multi-operand exponent alignment and parallel addition of mantissa.
[0049] The cooperation relationship between the accumulation operation module 130, the matrix multiplication and addition operation module 120, and the input module 110 is as follows: The exponent comparison submodule extracts the 10-bit exponent field of all floating-point elements from the aforementioned two inputs, and obtains the maximum exponent as the common exponent benchmark for the fifth operand through combinational logic comparison; the alignment and shift submodule performs right shift alignment on the corresponding mantissa based on the difference between the exponent of each operand and the maximum exponent, and generates three auxiliary bits: Guard, Round, and Sticky (GRS); the addition operation submodule performs 5-way parallel addition reduction on the aligned mantissa (including GRS), and outputs the mantissa field of the fifth operand, thereby completing the numerical fusion of the intermediate result and the initial value. The output fifth operand is a mantissa and exponent pair with a unified exponent and unnormalized FP32 format, providing the input basis for the normalization processing of the output module 140, ensuring that the BP16×BP16 intermediate result and the FP32 initial value are accurately added under the same exponent scale, avoiding the precision loss caused by exponent truncation.
[0050] The output module 140 is the final output unit for normalizing, rounding, overflow detection and format encapsulation of the execution result. Its input is connected to the fifth operand output by the accumulation module 130, and its output generates the target operand. The output module 140 is configured to output the target operand based on multiple fifth floating-point elements, wherein the target operand is the final result of the standard FP32 format (Sign(1) + Exp(8) + Man(23)) and conforms to the RISC-VF extension specification.
[0051] The output module 140 may include a leading zero detector (LZC), a normalization shifter, a rounding logic unit, and an exception flag generator. Its function is to convert the denormalized mantissa of the fifth operand into a normalized FP32 representation and perform rounding operations according to the five rounding modes (RNE / RTZ / RDN / RUP / RMM) defined by RISC-V.
[0052] Meanwhile, the connection between the output module 140 and the accumulation module 130 is as follows: the leading zero detector receives the mantissa of the fifth operand and outputs the number of leading zeros; the normalization shifter performs a left shift based on this number and adjusts the exponent synchronously; the rounding logic unit combines the GRS bit and the rounding mode to generate the Round bit and performs rounding addition; if rounding causes carry overflow, the secondary right normalization logic is triggered, the exponent is incremented by 1 and the mantissa is shifted right by 1 bit; the exception flag generator checks whether the final exponent exceeds the valid range of FP32 [1, 254], and sets the corresponding FCSR exception bit (OF / UF / NV) in cases of overflow / underflow / invalid operation, thereby completing the standardized encapsulation of numerical results and system-level exception feedback. The target operand of its output can be directly written back to the RISC-V vector register file or memory for subsequent instruction consumption.
[0053] In some embodiments, such as Figure 2 As shown, the BP16 matrix multiplication and addition circuit based on the RISC-V architecture also includes a preprocessing module 150; wherein, the input terminal of the preprocessing module 150 is connected to the input module 110, and the output terminal of the preprocessing module 150 is connected to the matrix multiplication and addition module 120; the preprocessing module 150 is configured to preprocess the first operand, the second operand, and the third operand respectively to obtain the preprocessed first operand, the second operand, and the third operand.
[0054] In this embodiment, the preprocessing module 150 can be understood as a logic unit that performs format parsing, semantic discrimination, and data normalization on the three floating-point operands input to the circuit before the matrix multiplication and addition operation is executed. It can convert the original BP16 / FP32 mixed format input into a data stream with consistent exponent and mantissa semantics and a clear normalization state that can be directly processed by the subsequent operation module. The preprocessing module 150, together with the input module 110 and the matrix multiplication and addition operation module 120 defined in this application, constitute a serial data path. Its input end receives the unprocessed original operands, and its output end provides the matrix multiplication and addition operation module 120 with intermediate operands after bit field separation, special value identification, and hidden bit recovery, thereby avoiding the special value judgment logic being embedded in the multiplication or accumulation sub-modules, reducing the complexity of the control path and hardware redundancy.
[0055] The input terminal of the preprocessing module 150 is connected to the input module 110 via a hard-wired connection implemented through a standard data bus, and its bus width is adapted to the parallel bit width of the first operand, the second operand, and the third operand.
[0056] For example, when the input module 110 periodically outputs 8 BP16 elements (128 bits in total) and the corresponding FP32 accumulated initial value (256 bits in total) with a 512-bit wide bus, the input interface of the preprocessing module 150 is configured to synchronously receive the multi-channel parallel data stream and complete the parallel splitting and preliminary discrimination of the three operands in a single cycle. This avoids the introduction of additional delay registers and only includes combinational logic and timing buffers to meet the timing constraints of RISC-V vector extension (V extension) or custom coprocessor interface.
[0057] The output of the preprocessing module 150 is connected to the matrix multiplication and addition module 120. It is an interface that uses an aligned extended format for output. Its output data format is strictly matched with the input protocol of the matrix multiplication and addition module 120. For example, the first and second operands after preprocessing are both extended to the format Sgn(1) + Exp(10) + Man(24). The third operand after preprocessing is kept in FP32 format, but its exponent field is uniformly biased to ensure that the three output data are comparable and computable in terms of exponent dynamic range, mantissa precision and sign representation. A first-level register group can be set between the output and the matrix multiplication and addition module 120 for cross-clock domain synchronization or pipeline partitioning, but without changing the data semantics.
[0058] The preprocessing module 150 is configured to preprocess the first operand, the second operand, and the third operand respectively. It can be three independent preprocessing channels that are executed in parallel and processed in a homogeneous manner. Each channel shares the same control logic and discrimination rules, but each maintains an independent status register and data path.
[0059] The preprocessing of the first and second operands focuses on BP16 format parsing and denormalized number normalization preparation, while the preprocessing of the third operand focuses on FP32 exponent consistency calibration and zero / Inf / NaN semantic alignment. All three preprocessing processes are driven by the same control signal to ensure timing coordination and avoid data mis-times caused by differences in processing progress.
[0060] The preprocessing module 150 is configured to obtain the first, second, and third operands after preprocessing, in order to output intermediate data with well-defined format boundaries: in the first and second operands after preprocessing, each floating-point element contains an explicit sign bit, a 10-bit extended exponent (including 2 bits for overflow reservation), and a 24-bit extended mantissa (including implicit bits and GRS bits), and all denormalized numbers have undergone LZC detection and left-shift preprocessing.
[0061] In the third operand after preprocessing, the exponent field of each floating-point element can be re-encoded according to the signed two's complement representation. Its mantissa field has been padded with hidden bits and extended to 24 bits to align with the mantissa width of the fourth operand. It does not constitute a new data type definition, but is an engineering intermediate representation based on the original BP16 / FP32 semantics. Its format can fluctuate in the range of Sgn(1)+Exp(9)+Man(23) to Sgn(1)+Exp(11)+Man(25) according to the actual circuit area and performance trade-off.
[0062] Specifically, the preprocessing module 150 operates as follows: At the beginning of each operation cycle, the input module 110 sends the first, second, and third operands of the current batch into the preprocessing module 150 in parallel. The preprocessing module 150 first performs bit field splitting on each floating-point element in each operand, separating the sign bit, exponent field, and mantissa field. Then, it calls a unified special data detection logic to determine whether the element is NaN, Inf, Zero, normalized, or denormalized based on the combination of exponent and mantissa. For denormalized numbers, a hidden bit recovery mechanism is activated to pad the hidden bit of the mantissa with 0 and trigger leading zero detection. For normalized numbers, the hidden bit is padded with 1. Finally, the sign, adjusted exponent, and extended mantissa of each operand are assembled into a unified intermediate format and output to the matrix multiplication and addition module 120. This allows the representation of the original numerical value to be reconstructed without changing its mathematical semantics, providing a uniform and clear input foundation for subsequent high-throughput matrix multiplication and addition.
[0063] As an example, the specific implementation of the scheme in this application is as follows: Taking (8×4)×(4×8) matrix multiplication and addition as an example, the input module 110 provides 8 BP16 elements (first operand), 8 BP16 elements (second operand), and 16 FP32 elements (third operand) to the first three operands in one cycle; the preprocessing module 150 synchronously completes the bit field splitting of all 64 floating-point elements in this cycle, identifies 2 non-normalized BP16 elements and 1 FP32 Inf value; performs LZC detection and left shift preprocessing on the non-normalized BP16 elements, and marks the FP32 Inf value with a special path identifier; then the three operands can be packaged into the format Sgn(1)+Exp(10)+Man(24) and output to the matrix multiplication and addition operation module 120; the matrix multiplication and addition operation module 120 performs row and column multiplication accordingly, and the entire preprocessing process takes no more than 1 cycle without introducing additional pipeline pauses.
[0064] In this application, by adding a preprocessing module 150 and performing format unification and numerical regularization on the three operands, the precision of the mantissa and the compatibility of the exponent dynamic range are ensured between the first and second operands during BP16×BP16 multiplication. At the same time, the preprocessing module 150 performs exponent remapping and mantissa expansion on the third operand, which enables the initial value of FP32 accumulation to be aligned with the intermediate result of BP16 multiplication output without error in the subsequent accumulation stage. In addition, the preprocessing module 150 has built-in special data detection and independent protection logic, which can intercept Denorm, NaN and extremely small exponent scenarios in advance to avoid them causing cascading anomalies in the main matrix multiplication and addition path, thereby improving the numerical robustness and operational predictability of the entire BP16 matrix multiplication and addition circuit.
[0065] In some embodiments, the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element are represented using signed two's complement.
[0066] In this embodiment, the exponent of the floating-point element in the third operand after preprocessing can be understood as the exponent value obtained after the preprocessing module 150 completes the format conversion, normalization and bias adjustment. The exponent value is represented in signed two's complement form with a bit width of not less than 10 bits, which can provide the subsequent exponent alignment module with a data format that can directly participate in arithmetic comparison and difference operation, avoiding the maximum value selection error caused by the distortion of negative exponent mapping under unsigned encoding. At the same time, the exponent and the exponent of the fourth floating-point element maintain the same format in the data path, and the two are jointly input to the exponent comparison module, forming a unified input basis for the exponent alignment operation.
[0067] In the preprocessed third operand, the exponent of the floating-point element is an integer value in two's complement format, with the highest bit being the sign bit and the remaining bits being the numeric bits. This does not change the mathematical meaning of the exponent of the original operand, but only its hardware encoding form. At the same time, it can also be converted to two's complement synchronously with the exponent of the fourth floating-point element in the circuit and sent to the exponent comparison module through the same bus structure. This allows the exponent comparison module to reuse the standard signed comparator unit without the need for additional sign judgment logic, thereby reducing the complexity of the control logic. In addition, this two's complement representation supports the direct calculation of negative exponent difference, which can provide a deterministic basis for the subsequent right shift alignment amount (i.e., exponent difference).
[0068] The exponent of the fourth floating-point element is the exponent part of the intermediate result generated by the matrix multiplication and addition module 120 after performing BP16×BP16 multiplication on the first operand and the second operand. This exponent is encoded in two's complement by the exponent adjustment module in the preprocessing module 150 before output.
[0069] In this application, the exponents of the floating-point elements in the third and fourth floating-point elements after preprocessing are both represented using signed two's complement. The exponent comparison module can directly implement maximum exponent filtering based on a standard binary comparator, avoiding the two-stage logic delay required by traditional bias representation, which requires debiasing before comparison. Two's complement representation provides a clear definition for negative exponents, providing a data basis for independent protection processing of the minimum negative value. Thus, when a two's complement value of -256 is detected, a dedicated protection path can be triggered to prevent normalization anomalies or invalid rounding caused by exponent underflow. Unified encoding eliminates the format conversion overhead in cross-precision exponent operations, improving the clock frequency and energy efficiency of the entire matrix multiplication-accumulation pipeline.
[0070] In some embodiments, when the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element reach the minimum negative value represented by the sign complement, independent protection processing is performed on the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element.
[0071] In this embodiment, the exponents of the floating-point elements in the third and fourth floating-point elements after preprocessing are represented using signed two's complement, with a bit width of not less than 10 bits. For example, it can be 10, 11, or 12 bits. The specific bit width can be set by balancing the precision requirements of BP16 operations in RISC-V vector extension with hardware resource constraints.
[0072] When the exponent of a floating-point element reaches the smallest negative value represented by the sign complement, it corresponds to the extreme form of all 1s followed by 0s in the complement encoding, such as -512 in 10-bit complement (i.e., binary 1000000000). Its value is usually subjected to conventional arithmetic operations such as negation, left shift (for normalization), addition (for exponent alignment offset), or maximum value filtering by the comparator. It is very easy to cause logical errors or undefined behavior due to overflow.
[0073] The minimum negative value is identified as a type of boundary exponent state requiring special handling in this application. Its functional meaning is to characterize the deepest underflow exponent level that the BP16 denormalized number may reach after preprocessing. It is named based on the fact that it constitutes a systematic risk source in the exponent alignment and normalization link.
[0074] Independent protection processing can be understood as setting up an extreme value detection and response module in the exponential data path that is independent of the main operation path. This module does not participate in regular arithmetic operations, but only performs pattern matching and condition replacement on the input exponential value.
[0075] When the exponents of the third and fourth floating-point elements in the currently processed operand reach the minimum negative value represented by the sign complement, an overflow occurs. This forms an action path with the exponent comparison module and the normalization module. That is, if the input contains the minimum negative value during the process of selecting the maximum exponent, the exponent comparison module must avoid misjudging it as the maximum value. If the input contains the minimum negative value during the left / right normalization adjustment, the normalization module must prevent the output of an illegal FP32 exponent due to exponent subtraction overflow. This ensures that even under the most unfavorable input combination (such as full Denorm input with multiple levels of accumulation), the controllability of the exponent field and the physical consistency of mantissa alignment can be maintained, thus forming a stable and predictable floating-point operation state flow.
[0076] During the independent protection process of the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element, when it is detected that the exponent of any floating-point element in the preprocessed third operand or the exponent of any floating-point element in the fourth floating-point element is equal to the smallest negative value in the signed two's complement representation, a dedicated protection path is activated. This protection path does not rely on general arithmetic logic units, but is implemented through hard-wired logic.
[0077] Specifically, in the exponent comparison module, the smallest negative value can be set as an invalid candidate, forcibly excluding it from participating in the maximum value competition; in the alignment shift module, when the smallest negative value is used as the minuend in the exponent difference calculation, a fixed safety offset (e.g., 0 or 1) can be clamped to output instead of performing actual subtraction; in the normalization module, when the smallest negative value is used as the input exponent in the left normal exponent correction, the exponent update can be frozen and a preset compensation value can be enabled (e.g., the exponent can be forcibly set to -128 to match the lower limit of the FP32 minimum positive normalized number exponent).
[0078] In addition, when the exponent comparison module selects the largest exponent from the exponents of the floating-point elements in the preprocessed third operand and the exponents of multiple fourth floating-point elements, if any input is the smallest negative value, the input is automatically masked by the logic gate circuit, and only the valid exponents participate in the comparison. Subsequently, the alignment and shift module calculates the right shift number of the mantissa of each fourth floating-point element based on the selected largest exponent. At this time, if the exponent of a certain fourth floating-point element was originally the smallest negative value, its right shift number is forced to be set to the maximum allowed right shift amount, and its mantissa high bits are padded with the sign bit to maintain the numerical meaning. Thus, even under extreme underflow input conditions, the accumulation operation can still generate an intermediate result with physical interpretability, rather than triggering an abnormal interrupt or uncontrollable output.
[0079] In this application, a protection logic independent of the general arithmetic path is enabled when the minimum negative value occurs, avoiding alignment failure caused by exponent inversion or subtraction overflow. At the same time, the protection logic can form a closed loop with the exponent comparison, alignment shift, and normalization modules, thereby ensuring the numerical robustness of BP16 matrix multiplication and addition operations in the entire input domain. In addition, the protection action can be implemented based on hardware hardwired without software intervention, without introducing additional timing overhead, and maintaining the single-cycle issue characteristic of RISC-V vector instructions.
[0080] In some embodiments, the preprocessing module 150 includes a splitting module, a special data detection module, and a hidden bit recovery module; wherein, the splitting module is connected to the input module 110, the hidden bit recovery module, the matrix multiplication and addition module 120, and the accumulation module 130, respectively; the special data detection module is connected to the input module 110, the hidden bit recovery module, and the accumulation module 130, respectively; and the hidden bit recovery module is connected to the matrix multiplication and addition module 120 and the accumulation module 130, respectively; the splitting module is configured to split the first floating-point element, the second floating-point element, and the third floating-point element into bits respectively to obtain the sign bit, exponent, and mantissa in the first floating-point element, the second floating-point element, and the third floating-point element; the special data detection module is configured to perform normalization detection on the mantissa in the first floating-point element, the second floating-point element, and the third floating-point element respectively to obtain the normalization detection result; and the hidden bit recovery module is configured to recover the hidden bit of the mantissa in the first floating-point element, the second floating-point element, and the third floating-point element based on the normalization detection result.
[0081] In this embodiment, the splitting module can be understood as a logic circuit that performs bit-field separation on the input BP16 format floating-point element or FP32 format floating-point element. Its input is 16-bit BP16 data (Sign(1) + Exp(8) + Man(7)) or 32-bit FP32 data (Sign(1) + Exp(8) + Man(23)), and its output includes independent sign bit signal, exponent bit signal and mantissa bit signal.
[0082] The splitting module can be implemented using a combination of multiplexers and fixed shifters, or it can be constructed by decoding logic in conjunction with a register group. The splitting module can provide the original bit field decoupling output for each subsequent sub-module without changing the meaning of the values, and only completes the physical depackaging of the data format.
[0083] In this application, the splitting module can work in parallel with the special data detection module and the hidden bit recovery module to ensure the timing predictability of the preprocessing process and the low coupling between pipeline stages. The sign bit, exponent bit and mantissa bit output by the splitting module are sent to the corresponding downstream modules for further discrimination and processing, thereby supporting the modular division of labor in the entire preprocessing link.
[0084] The special data detection module can be understood as a combinational logic circuit that classifies and identifies the combination of mantissa and exponent of floating-point elements according to the IEEE 754 and RISC-V F extension specifications. Its input is the mantissa and exponent bits output by the splitting module, and its output is one or more bits of normalization detection result, which is used to characterize whether the current floating-point element belongs to a normalized number, a denormalized number, zero, infinity or NaN.
[0085] The special data detection module can be implemented based on the joint judgment of the exponent value and the mantissa value: when the exponent is 0 and the mantissa is 0, it is judged as zero; when the exponent is 255 and the mantissa is 0, it is judged as Inf; when the exponent is 255 and the mantissa is not 0, it is judged as NaN; when the exponent is 0 and the mantissa is not 0, it is judged as a denormalized number; when 1≤exponent≤254, it is judged as a normalized number.
[0086] The special data detection module can serve as the central hub for type discrimination in the preprocessing flow. Its detection results directly drive the action selection of the hidden bit recovery module and provide effective input enable control for the leading zero detection module. It also avoids inserting complex branch judgments in the multiplication or addition path, thereby improving the timing robustness and frequency achievability of the overall circuit.
[0087] The hidden bit recovery module can be understood as a logic circuit that dynamically completes the hidden highest bit of the split mantissa based on the normalized detection result output by the special data detection module. For normalized numbers, the module inserts 1s into the high bits of the mantissa to form a standard mantissa in the format of 1.xxxx. For denormalized numbers, the module inserts 0s into the high bits of the mantissa to retain the format of 0.xxxx.
[0088] The hidden bit recovery module can be implemented by a multiplexer with constant drive, or it can be integrated into the write path of the mantissa register. Its inputs include the original mantissa output by the split module and the type flag output by the special data detection module. The output is the extended mantissa with hidden bits.
[0089] In this application, the module is positioned to unify the mantissa representation and provide standardized input for subsequent BP16×BP16 multiplication and FP32 accumulation operations; it forms a strongly bound closed loop with the special data detection module: the detection result determines the recovery method, and the recovery result verifies the correctness of the detection.
[0090] The hidden bit recovery module outputs the mantissa with hidden bits, which is synchronously sent to the multiplication module and the accumulation module 130 as a unified operation object for subsequent exponent alignment, mantissa alignment and parallel reduction.
[0091] Specifically, the splitting module, special data detection module, and hidden bit recovery module operate in parallel during the preprocessing stage: the splitting module completes bit field unpacking of the BP16 / FP32 word in the first cycle; the special data detection module completes type discrimination based on the original mantissa and exponent in the second cycle and outputs the detection result; the hidden bit recovery module completes the insertion of the hidden bit of the mantissa based on the detection result in the third cycle and outputs the normalized mantissa; there is no data dependency blocking among the three, only a control signal synchronization relationship; this parallel structure makes the preprocessing delay stable at 3 cycles, and each module can independently optimize the layout and routing, significantly improving the synthesis timing convergence; at the same time, the module interface is clearly defined (only transmitting bit field signals and type flags), supporting cross-process node reuse and IP packaging.
[0092] As an example, in a RISC-V vector extension instruction set environment, when executing a BP16 matrix multiplication and addition instruction of type vfmacc.vv, the input module 110 sends the first BP16 operand (e.g., vs1), the second BP16 operand (e.g., vs2), and the third FP32 operand (e.g., vd) from the vector register group to the preprocessing module 150 in parallel; the splitting module performs a 1→8→7 bit split on each 16-bit BP16 element, outputting three sets of signals S1 / E1 / M1; it performs a 1→8→23 bit split on each 32-bit FP32 element, outputting three sets of signals S3 / E3 / M3; the special data detection module receives M1 / E1 and M3 / E3 respectively... Determine whether each BP16 element in vs1 / vs2 is adenorm / norm / zero / Inf / NaN, and whether each FP32 element in vd is adenorm / norm / zero / Inf / NaN, and output the corresponding 2-bit type code; the hidden bit recovery module inserts 1 or 0 into the high bits of M1 and M3 respectively according to the code, generating extended mantissas M1_ext (8 bits) and M3_ext (24 bits) for subsequent multiplication and accumulation modules to call; the whole process completes the parallel splitting and detection of all elements in a single cycle, and outputs the normalized mantissa with hidden bits in the next cycle, satisfying the cycle time constraint of the four-cycle matrix multiplication and addition pipeline.
[0093] In this application, the splitting module, the special data detection module, and the hidden bit recovery module constitute a three-level pipelined preprocessing link. The splitting module first completes the physical decoupling of S / E / M of BP16 elements. The special data detection module determines its semantic type based on the joint E and M. The hidden bit recovery module dynamically injects hidden bits according to the type code and outputs the normalized mantissa, so that the same hardware circuit can process the BP16 format without ambiguity and provide structurally consistent, semantically clear, and numerically complete input data for the subsequent matrix multiplication and addition operation module 120 and accumulation operation module 130. This link does not change the meaning of the original data values, but only completes the format parsing and numerical explicitation. All operations are completed in a single cycle and do not affect the overall pipeline depth.
[0094] In some embodiments, the preprocessing module 150 further includes a leading zero detection module, a mantissa normalization module, and an exponent adjustment module; wherein, the leading zero detection module is connected to the matrix multiplication and addition module 120, the accumulation module 130, the hidden bit recovery module, and the mantissa normalization module, respectively; the mantissa normalization module is connected to the matrix multiplication and addition module 120 and the accumulation module 130, respectively; and the exponent adjustment module is connected to the splitting module, the leading zero detection module, and the matrix multiplication and addition module 120, respectively; the leading zero detection module is configured to perform leading zero detection on the unnormalized mantissas in the first floating-point element, the second floating-point element, and the third floating-point element, respectively, to obtain a first leading zero detection result; the mantissa normalization module is configured to normalize the mantissas in the first floating-point element and the second floating-point element based on the first leading zero detection result; and both the exponent adjustment module and the mantissa normalization module are configured to process the exponent and the unnormalized mantissa in the third floating-point element based on the first leading zero detection result, respectively, to obtain the preprocessed third operand.
[0095] In this embodiment, the leading zero detection module can be understood as a logic unit used to perform a leading zero count (LZC) operation on the denormalized mantissa of the input floating-point element. Its output is a first leading zero detection result representing the number of consecutive zeros in the high-order bits of the denormalized mantissa. The first leading zero detection result is an unsigned integer, and the bit width is set according to the minimum supported denormalized mantissa length. For example, in the BP16 format, the mantissa is 7 bits, and the corresponding LZC output bit width is 3 bits (which can represent 0-7).
[0096] The leading zero detection module is responsible for the identification and quantization of denormalized numbers, and its detection results serve as the control basis for subsequent normalization shift and exponential compensation.
[0097] Meanwhile, the leading zero detection module is connected to the splitting module, hidden bit recovery module, mantissa normalization module, and matrix multiplication and addition module 120 via a multi-channel data bus and control signal line. The first leading zero detection result output by the module is synchronously distributed to the mantissa normalization module and the exponent adjustment module to achieve coordinated response of the three floating-point element processing paths. This allows the denormalized mantissa to complete the valid bit positioning before entering the multiplication operation, avoiding invalid low bits from participating in the operation, thereby reducing power consumption and improving computational efficiency.
[0098] The mantissa normalization module can be understood as performing a left shift alignment operation on the denormalized mantissa based on the first leading zero detection result, and simultaneously updating the combinational logic circuit of the implicit bits of the corresponding floating-point element and the mantissa precision representation.
[0099] The mantissa normalization module can be a pure combinational logic structure or it can include a first-level register for timing constraint optimization. Its inputs include the original mantissa from the splitting module, the normalization status flag from the special data detection module, and the first leading zero detection result from the leading zero detection module. Its output is the normalized mantissa and the corresponding normalization offset.
[0100] The mantissa normalization module can be connected to the leading zero detection module, the matrix multiplication and addition module 120, and the accumulation module 130 via a width-matched data path.
[0101] In this application, the mantissa normalization module is used to normalize the unnormalized mantissas in the first and second floating-point elements, converting them into a normalized representation in the form of 1.xxxx, and explicitly restoring the hidden bits. It works in conjunction with the leading zero detection module to form the core execution unit for the Denorm to Norm conversion, enabling BP16 multiplication operations to be performed under the premise of unified normalization, ensuring the consistency of the product mantissa precision and the predictability of the exponent calculation.
[0102] The exponent adjustment module can be understood as an arithmetic logic unit that performs compensation correction on the exponent field of the third floating-point element based on the first leading zero detection result. The exponent adjustment module can be an adder or a lookup table structure, and its inputs include: the original exponent from the splitting module and the first leading zero detection result from the leading zero detection module; its output is the compensated exponent value.
[0103] The exponent adjustment module can be connected to the splitting module, the leading zero detection module, and the matrix multiplication and addition module 120 via a bus with matching exponent bit width.
[0104] In this application, the exponent adjustment module and the mantissa normalization module work together. For the unnormalized number in the third floating-point element (i.e., the accumulated input C), while normalizing its mantissa, the exponent is adjusted simultaneously to ensure that the exponent reference of the third floating-point element is consistent with the exponent reference of the first and second floating-point elements after normalization. The exponent compensation amount performed by the exponent adjustment module is equal to the number of left shifts indicated by the first leading zero detection result. This ensures that the subsequent exponent alignment module uses the exponent value under the same normalization level when comparing the exponents of multiple floating-point elements, eliminating the exponent alignment deviation caused by the third floating-point element not being synchronously normalized, and improving the accuracy and robustness of numerical alignment in the accumulation stage.
[0105] Specifically, when the input first, second, or third floating-point element is determined to be a denormalized number by the special data detection module, the leading zero detection module immediately performs a leading zero count on its mantissa and outputs the first leading zero detection result. The first leading zero detection result can drive the mantissa normalization module to perform a left shift LZC bit operation on the mantissa of the first and second floating-point elements, and fill the low bits with 0, while restoring the implicit bits to form a complete normalized mantissa. On the other hand, the first leading zero detection result can be synchronously sent to the exponent adjustment module to subtract LZC from the original exponent of the third floating-point element to offset the exponent reduction effect caused by the left shift of the mantissa, thereby maintaining numerical identity. This ensures that all input floating-point elements are in a uniform normalized state and their exponents are comparable before entering the matrix multiplication and addition operation module 120, thus laying the foundation for subsequent high-precision accumulation alignment.
[0106] As an example, when input module 110 receives a set of first and second operands in BP16 format and a third operand in FP32 format, the splitting module first unpacks each operand into a sign bit, exponent, and mantissa field; the special data detection module determines that a certain third floating-point element is a denormalized number (i.e., the exponent is 0 and the mantissa is not 0), and sends the determination result to the hidden bit recovery module and the leading zero detection module; the leading zero detection module performs LZC operation on the 7-bit mantissa of the third floating-point element, outputting lzc_denorm=3; the mantissa normalization module then... The mantissas of the first and second floating-point elements are shifted left by 3 bits, changing their format from 0.001xxxx to 1.xxx000. Simultaneously, the exponent adjustment module subtracts 3 from the FP32 exponent of the third floating-point element, resulting in Exp_out = -3. Since the FP32 exponent is represented in signed two's complement and must meet format constraints, this negative exponent participates in the calculation through the maximum exponent selection stage in the subsequent alignment and reduction process. Ultimately, it, together with the extended exponent of the multiplication result, determines the accumulation alignment benchmark, avoiding alignment errors or underflow misjudgments caused by uncorrected denorm.
[0107] In this application, by setting a leading zero detection module, the valid starting position of the denormalized mantissa can be accurately identified; the mantissa normalization module performs left shift normalization on the mantissa of the first floating-point element and the second floating-point element based on the detection result, which can eliminate redundant low-bit operations when the denormalized number participates in multiplication and improve energy efficiency; the exponent adjustment module simultaneously performs compensation correction on the exponent of the third floating-point element, which can ensure the consistency of the three input data on the exponent reference, ensure the atomicity and timing determinism of the Denorm processing, avoid cross-module asynchronous errors, and thus support the high reliability and low power consumption implementation of BP16 to FP32 matrix multiplication and addition operations under the RISC-V architecture.
[0108] In some embodiments, the accumulation module 130 includes an exponent comparison module, an alignment shift module, and an addition module; wherein, the exponent comparison module is connected to the preprocessing module 150, the alignment shift module, and the output module 140, and the addition module is connected to the preprocessing module 150, the alignment shift module, and the output module 140, respectively; the exponent comparison module is configured to select the largest exponent from the exponents of the floating-point elements in the preprocessed third operand and the exponents of multiple fourth floating-point elements to obtain the exponent of the fifth floating-point element; the alignment shift module is configured to right-shift and align the mantissas of multiple fourth floating-point elements based on the largest exponent to obtain multiple right-shifted and aligned mantissas; the addition module is configured to perform an addition operation on the mantissas of the floating-point elements in the preprocessed third operand and the multiple right-shifted and aligned mantissas to obtain the mantissa of the fifth floating-point element.
[0109] In this embodiment, the exponent comparison module is a logic circuit unit used to perform parallel comparisons between the exponents of multiple input floating-point elements and output the maximum value. Its inputs include: the exponents of each floating-point element in the third operand output by the preprocessing module 150 (i.e., the exponents of the preprocessed FP32 format accumulated initial value), and the exponents of multiple fourth floating-point elements output by the matrix multiplication and addition module 120 (i.e., the intermediate FP32 precision exponents formed after the FP8×FP8 multiplication result is expanded).
[0110] The exponent comparison module uses a chain of bit-by-bit comparators or a tree-structured comparison module to filter the maximum value of multiple exponents. Its output is a set of exponent values that correspond one-to-one with the fifth floating-point element. These exponent values are synchronously distributed to the alignment shift module and the output module 140 and can be used as the reference for subsequent normalization and rounding.
[0111] The exponent comparison module can establish a unified exponent alignment benchmark, avoiding the loss of significant mantissa bits due to exponent dispersion. It and the alignment shift module form the starting link of the exponent-mantissa collaborative path. By outputting the maximum exponent, it drives the calculation of the right shift of all mantissa bits to be accumulated, thereby ensuring that fixed-point addition is completed on the same order of magnitude for multiple data.
[0112] The alignment shift module is a hardware shift unit used to perform an unsigned right shift and zero-padding / sign-padding of the mantissas of each path based on the maximum exponent. Its inputs include: the mantissas of each floating-point element in the preprocessed third operand, the mantissas of multiple fourth floating-point elements, and the maximum exponent output by the exponent comparison module.
[0113] The alignment and shift module first calculates the difference between the exponent corresponding to the mantissa of each path and the maximum exponent. This difference is the number of bits to be shifted right. For mantissa paths with a difference greater than 0, the corresponding number of bits is shifted right, and zeros are added to the high bits (for positive numbers) or a sign bit is added (for negative numbers, an arithmetic right shift is used). For paths with a difference equal to 0, the mantissa is output as is.
[0114] In addition, the alignment shift module is configured to synchronously generate a guard bit, a round bit, and a sticky bit during the right shift process. The sticky bit is the logical OR result of all the low-order bits that are shifted out.
[0115] The alignment and shift module can achieve lossless alignment of multi-precision floating-point data in the fixed-point domain. It works with the exponent comparison module to form an alignment decision-alignment execution closed loop, and together with the addition module, it forms the backbone of the alignment summation data flow. This allows the mantissas of the 8-way fourth floating-point elements and the mantissas of the 1-way third floating-point elements to participate in addition under the same exponent benchmark, ensuring that all valid information enters the subsequent reduction process.
[0116] The addition module is a tree-structured adder used to perform parallel addition of multiple fixed-point mantissas. Its inputs include the mantissas of each floating-point element in the preprocessed third operand (already aligned to the maximum exponent) and the mantissas of multiple right-shifted aligned fourth floating-point elements. The addition module can complete the core numerical fusion of the accumulation stage. Its input comes directly from the output of the alignment and shift module, and the output is sent to the output module 140 as the mantissa of the fifth floating-point element.
[0117] The addition module can be implemented using a multi-level parallel addition tree, such as a Wallace tree or Dadda tree structure formed by cascading a 4-2 compressor and a 3-2 compressor. It can reduce the 5-way wide-bit mantissa to two-way (sum + carry) within 2-3 clock cycles, and then obtain a single-way complete mantissa result through the final adder. The mantissa of the multiple fifth floating-point elements output by this module is a set of wide-bit mantissas (e.g., 28 bits) after reduction, which retains enough high bits to accommodate carry overflow and provides the input basis for subsequent normalization and rounding. This can replace the traditional serial accumulation chain, significantly improving throughput and energy efficiency. At the same time, the addition module and the alignment shift module form a data consumption relationship, receiving the aligned mantissa output by the addition module. In addition, its output is directly connected to the leading zero detection and normalization processing unit in the output module 140, forming a continuous data flow closed loop.
[0118] Specifically, the addition module, together with the exponent comparison module and the alignment shift module, constitutes the complete data path of the accumulation module 130, realizing a deterministic pipeline of exponent alignment, mantissa alignment, and multi-way summation, providing high-fidelity mantissa input for subsequent normalization.
[0119] In this application, the accumulation operation process is as follows: The exponent comparison module first determines the maximum exponent value of the current accumulation group from the exponent of the third operand (FP32 format, 8 bits) and the exponent of the fourth operand (extended exponent obtained by multiplying BP16×BP16, 10 bits) by comparing bit by bit and prioritizing encoding. The maximum exponent is broadcast to the alignment shift module, which calculates the number of bits that the mantissa of each fourth floating-point element needs to be shifted to the right and performs a shift operation with GRS padding. At the same time, the mantissa of the third operand is not shifted (because its exponent has been confirmed as the maximum or aligned) and directly enters the addition operation module. The addition operation module synchronously inputs the unshifted mantissa of the first channel and the shifted mantissa of the fourth channel into the parallel addition tree, completes the accumulation of the mantissa of the five channels in a single cycle or two-stage pipeline, and outputs a set of wide-bit mantissa results. The bit width of the result is sufficient to cover all possible carry, and the decimal point position strictly corresponds to the maximum exponent, providing a structurally complete input condition for subsequent normalization processing.
[0120] As an example, in the (8×4)×(4×8) to (8×8) matrix multiplication and addition scenario, each fifth floating-point element corresponds to the sum of 4 BP16 multiplication results (i.e., 4 fourth floating-point elements) and 1 FP32 accumulation initial value (i.e., the corresponding element in the third operand); the exponent comparison module performs parallel comparison of the exponents of these 5 operands. For example, when the exponent of the third operand is 132 and the exponents of the four fourth operands are 129, 131, 128, and 132 respectively, the module outputs the maximum exponent of 132; the alignment and shifting module adjusts the exponent accordingly. The mantissas of 129, 131, and 128 are right-shifted by 3 bits, 1 bit, and 4 bits respectively, and the GRS bits are filled. The mantissas of the two exponents of 132 (the mantissa of the third operand and the mantissa of a fourth floating-point element) are kept in their original positions. The addition module inputs the five aligned mantissas (all 27 bits) into a 4-2 compressor tree. After two stages of compression, two 28-bit signals, sum and carry, are generated. The final adder then produces a 28-bit fifth floating-point element mantissa. The high bits of this mantissa may contain redundant 1s, but the overall result meets the requirements of LZC detection and left-alignment.
[0121] In this application, the exponent comparison module uniformly selects the largest exponent, avoiding the truncation of significant bits due to excessive differences in exponents among multiple floating-point data; the alignment and shift module performs precise displacement based on the largest exponent and generates GRS bits, which can provide complete information support for subsequent rounding; the addition module adopts a tree structure to perform parallel reduction of the mantissas of 5 channels, completing all accumulation operations in a single cycle, meeting the high throughput timing constraints of the RISC-V vector unit, and ensuring the numerical integrity, precision controllability, and hardware execution efficiency of BP16 matrix multiplication and addition in the accumulation stage.
[0122] In some embodiments, the output module 140 is configured to perform leading zero detection on the mantissas of a plurality of fifth floating-point elements to obtain a second leading zero detection result, and to normalize the mantissas of the plurality of fifth floating-point elements based on the second leading zero detection result to generate the mantissa in the target operand; to adjust the exponents of the plurality of fifth floating-point elements based on the second leading zero detection result to generate the exponent in the target operand; and to output the target operand.
[0123] In this embodiment, the output module 140 can be used to perform final formatting on the fifth floating-point element formed after the accumulation operation. The output module 140 can convert the denormalized mantissa into the 1.xxxx normalized form required by the IEEE 754 FP32 standard and simultaneously correct the corresponding exponent, thereby ensuring that the output result meets the RISC-V floating-point specification requirements in terms of numerical representation, rounding behavior and exception flag generation.
[0124] The output module 140 works in conjunction with the exponent comparison module to receive the maximum exponent value of its output as a normalization reference. At the same time, the output module 140 can also work in conjunction with the addition module to receive the unnormalized mantissa and sign bit of its output, forming a complete input data stream. This enables integrated processing of mantissa left / right shift, exponent increment / decrement, GRS bit extraction and rounding decision, forming a closed-loop path from the accumulation result to the standard FP32 output.
[0125] Specifically, in the process of performing leading zero detection on the mantissa of multiple fifth floating-point elements to obtain the second leading zero detection result, this application can use a leading zero counter (LZC) circuit to perform parallel leading zero statistics on the mantissa of each fifth floating-point element output by the addition module to output the second leading zero detection result, which represents the number of consecutive zeros starting from the most significant bit.
[0126] The second leading zero detection result can be an integer between 0 and 23, where 0 indicates that the highest bit of the mantissa is 1 (normalized), and 23 indicates that the mantissa is all 0 (i.e., the result is zero).
[0127] Meanwhile, in the process of performing leading zero detection on the mantissas of multiple fifth floating-point elements to obtain the second leading zero detection result, this application can be executed based on the original code mantissa, or the mantissa in two's complement can be uniformly converted to the original code form after sign extension before execution; the second leading zero detection result can be directly used for subsequent normalized shift amount determination and GRS bit generation.
[0128] For example: when the second leading zero detection result is n, shifting left by n bits can align the most significant bit to the 23rd bit (the most implicit bit of the FP32 mantissa); when the second leading zero detection result is 0 and the two most significant bits of the mantissa are 10, it indicates that there is a carry overflow, and it is necessary to right-align by 1 bit and increment the exponent by 1.
[0129] In the process of normalizing the mantissas of multiple fifth floating-point elements based on the second leading zero detection result to generate the mantissas in the target operand, the programmable shifter can be controlled to perform left or right shift operations on the mantissas according to the second leading zero detection result.
[0130] When the second leading zero detection result n>0, a left shift of n bits is performed, with low-order bits padded with 0; when the second leading zero detection result n=0 and the high-order bit of the mantissa is 10, a right shift of 1 bit is performed, with high-order bits discarded and low-order bits padded with rounding bits; the shifted mantissa retains 23 significant decimal places and reserves 1 implicit bit space; this normalization process can be completed in conjunction with rounding operations, for example, during the left shift, low-order redundant bits are truncated simultaneously and three rounding auxiliary bits G / R / S are generated; or after the shift is completed, an independent rounding module generates the final Rnd bit based on the GRS bit and RISC-V rounding mode (RNE / RTZ / RDN / RUP / RMM) and performs addition; the normalized mantissa is a standard FP32 format 23-bit explicit mantissa, with its most significant bit always 1 (implicit) or all zeros (zero value output).
[0131] In the process of adjusting the exponents of multiple fifth floating-point elements based on the second leading zero detection result to generate the exponent in the target operand, the exponents of multiple fifth floating-point elements (denoted as Exp_max) output by the exponent comparison module can be arithmetically operated with the second leading zero detection result n.
[0132] When left-shifting by n bits, the exponent is updated to Exp_out = Exp_max - n; when right-shifting by 1 bit, the exponent is updated to Exp_out = Exp_max + 1. This exponent adjustment process is implemented using a signed adder, supporting a 10-bit signed exponent operation range [-162, 386]. The adjusted exponent is mapped to the FP32 exponent field [1, 254] after being limited: if Exp_out > 254, it is determined to be an overflow (OF), and Inf or a saturation value is output; if Exp_out < 1 and the mantissa is not zero, it is determined to be an underflow (UF), and a denormalized number or zero is output. This mapping process is completed synchronously with the exception flag generation logic.
[0133] Finally, during the output of the target operand, the normalized mantissa, adjusted exponent, and original sign bit can be packaged and output in FP32 format. The sign bit comes directly from the sign bit output by the addition module, the exponent is an 8-bit unsigned integer (bias 127), and the mantissa is a 23-bit explicit decimal. The output data is written to the vector register file in cycles via a 512-bit bus, with 8 FP32 elements output per cycle, for a total of 4 cycles to complete the output of 64 FP32 results. At the same time, the output process is controlled by the RISC-V vector instruction vfmopma.vv, supporting mask writing and broadcast modes.
[0134] The output module 140 operates as follows: it receives the unnormalized mantissa and sign bit from the addition module and the maximum exponent value from the exponent comparison module; it calls the LZC circuit to perform leading zero detection on each mantissa to obtain the second leading zero detection result; based on the result, it determines whether to shift left, shift right, or keep it unchanged, and drives the shifter to perform the corresponding operation; synchronously, it sends the second leading zero detection result to the exponent adjustment unit to calculate the output exponent; at the same time, during or after the shift, it generates three bits (G / R / S) based on the low-order truncation information of the mantissa and sends them to the rounding module to participate in the Rnd bit calculation; finally, it combines the sign bit, the adjusted exponent, the normalized mantissa, and the exception flag into a standard FP32 word, which is then driven to the external interface via the output buffer.
[0135] As an example, in the (8×4)×(4×8)→(8×8) matrix multiplication and addition operation, the accumulation operation module 130 outputs 64 fifth floating-point elements, each containing a 24-bit mantissa and a 10-bit exponent; the output module 140 receives all 64 mantissas and sends them to 64 parallel LZC units respectively to obtain 64 second leading zero detection results; for a certain fifth floating-point element, its mantissa is 0000_0000_1011_0011_0101_0000 (24 bits), and the LZC detects the second leading zero detection result. The result is 8; the normalization shift unit shifts the mantissa left by 8 bits to obtain 1011_0011_0101_0000_0000_0000 (high bits are padded with 0, low bits are discarded), and then the high 23 bits are truncated as the target mantissa; the exponent adjustment unit subtracts 8 from the original exponent (e.g., 135) to obtain 127, which is used as the target exponent; finally, the sign bit is concatenated with 0, the exponent is 127, and the mantissa is concatenated to output the target operand in standard FP32 format; the remaining 63 elements are processed in parallel according to the same process, and the overall output delay is stable and controllable, and does not change with the distribution of input data.
[0136] In this application, the output module 140 performs normalization processing on the mantissa of the fifth floating-point element based on the second leading zero detection result, thereby restoring the denormalized mantissa to the standard form of 1.xxxx; the corresponding exponent is adjusted synchronously to ensure that the numerical precision is not lost due to shift; the normalization process is deeply coupled with GRS bit generation and rounding mode selection, and can support all five rounding behaviors defined by RISC-V; the exponent adjustment range covers [-162, 386] and overflow / underflow criteria are set, which can accurately generate abnormal flags such as OF / UF / NX, so that the final state normalized output of the BP16 matrix multiplication and addition link can be completed without introducing a new structure.
[0137] In some embodiments, the matrix multiplication and addition module 120 includes a multiplexer and a multiplication module; wherein the multiplexer is connected to the input module 110 and the multiplication module, and the multiplication module is connected to the accumulation module 130; the multiplexer is configured to periodically output a plurality of first floating-point elements in a preset number of rows of the first operand, and the multiplication module is configured to perform row and column multiplication operations on the plurality of first floating-point elements in the preset number of rows of the first operand and the second floating-point elements in the second operand in the form of a matrix to obtain a plurality of fourth floating-point elements.
[0138] In this embodiment, the multiplexer can be understood as a hardware selection circuit for periodically scheduling the input data stream. Its input terminal receives the complete first operand from the input module 110 (e.g., an 8×4 matrix with a width of 512 bits and 32 BP16 elements, or a 4×8 matrix with a width of 512 bits and 32 BP16 elements), and its output terminal supplies a preset number of BP16 elements to the multiplication module in batches according to the instruction cycle.
[0139] The preset number of rows is 1 or 2. For example, in the (8×4)×(4×8) matrix calculation scenario, 2 rows × 8 columns of 16 first floating-point elements are output in each cycle, corresponding to the i-th row and i+1-th row of the multiplicand matrix.
[0140] The multiplexer is implemented using a static decoding structure or a dynamic timing control structure. Its selection logic is driven by the cycle count signal of the RISC-V vector instruction vwmacc.vv.
[0141] In this application, the multiplexer can realize the spatiotemporal multiplexing and scheduling of input data. Together with the input module 110, it forms a periodic fragmented supply path for BP16 data. It works with the multiplication module to complete the directional pairing of row and column dimensions, thereby avoiding wiring congestion and power consumption increase caused by full matrix broadcasting, and supporting the 4-cycle pipelined matrix multiplication and addition execution process.
[0142] The multiplication module can be a parallel multiplication array that supports BP16×BP16 fixed-point multiplication. Its internal structure includes a sign bit XOR gate, an exponent adder, a mantissa Booth-4 encoded multiplier, and implicit bit recovery logic.
[0143] The multiplication module can receive the first floating-point element of the preset row number (e.g., 8 BP16 elements in 2 rows × 4 columns) output by the multiplexer and the second floating-point element of the corresponding column number in the second operand (e.g., 32 BP16 elements in 4 columns × 8 rows). It is organized into multiple parallel multiplication paths (DP) according to the matrix row and column rules. Each path performs a BP16 × BP16 multiplication operation once and outputs a fourth floating-point element. The format of each fourth floating-point element is Sgn(1) + Exp(10) + Man(24), where the 10-bit exponent is generated by adding two BP16 exponents and subtracting the bias 127 and the leading zero compensation value, and the 24-bit mantissa is obtained by normalizing the left shift and expanding the padding of the 8-bit × 8-bit Booth-4 multiplication result. This multiplication module, together with the multiplexer, forms the first stage of the scheduling-computation two-stage pipeline. It is connected to the floating-point parallel reduction module by a full-width data bus to ensure lossless transmission of intermediate products.
[0144] Specifically, the multiplexer schedules the first operand in a four-cycle pipeline rhythm: the first cycle outputs rows 0-1, the second cycle outputs rows 2-3, the third cycle outputs rows 4-5, and the fourth cycle outputs rows 6-7.
[0145] Within each cycle, the multiplication module performs 8×4=32 parallel BP16×BP16 multiplications on the selected 2 rows of first floating-point elements (8 in total) and the corresponding 32 second floating-point elements in the 4 columns of the second operand, generating 64 fourth floating-point elements, which are organized into 16 groups (4 fourth floating-point elements in each group) by column. After four cycles, a total of 8 rows × 8 columns array is output, with each group containing 4 fourth floating-point elements, forming an intermediate result matrix of (8×4)×(4×8) matrix multiplication.
[0146] As an example, the specific implementation of the scheme in this application is as follows: Taking the (8×4)×(4×8) matrix multiplication and addition task as an example, the input module 110 provides the first operand A (8 rows × 4 columns BP16) and the second operand B (4 rows × 8 columns BP16) to the multiplexer; the multiplexer selects the 0-1 rows of A (a total of 8 BP16 elements) in the first cycle and sends them to the multiplication operation module; the multiplication operation module performs 8×4=32-way multiplication with the 0th column of B (4 BP16 elements) respectively to obtain 64 sixth floating-point elements, and groups them according to the column index of B (4 BP16 elements in each column), outputting 16 groups of arrays (each group of arrays includes 4 fourth floating-point elements), that is, the 8 arrays in the 0th and 1st rows of the result matrix C; and so on, after four cycles, all 8 rows are scheduled, and finally the 8 rows × 8 columns of arrays are output, each group of arrays includes 4 fourth floating-point elements, forming the intermediate result matrix C of (8×4)×(4×8) matrix multiplication.
[0147] In some embodiments, the matrix multiplication and addition module 120 further includes a processing module; wherein the processing module is an accumulation module 130; the multiplication module is configured to perform row and column multiplication operations on a plurality of first floating-point elements in a preset number of rows in the first operand and second floating-point elements in the second operand in the form of a matrix to obtain a plurality of sixth floating-point elements; the processing module is configured to extend the mantissas of the plurality of sixth floating-point elements and perform an alignment operation on the exponents of the plurality of sixth floating-point elements to obtain a fourth operand.
[0148] In this embodiment, the sixth floating-point element can be an intermediate result generated by the multiplication module after performing a matrix multiplication operation on multiple first floating-point elements in a preset number of rows in the first operand and second floating-point elements in the second operand.
[0149] The sixth floating-point element can serve as a transition point from BP16 multiplication precision to FP32 accumulation precision. It forms a format adaptation front-end with the processing module and achieves low-latency data flow through a direct signal path. As a result, the numerical integrity of the sixth floating-point element can be calibrated before entering the accumulation stage, thereby avoiding precision loss caused by mantissa truncation or exponent misalignment.
[0150] The processing module is configured to extend the mantissa of multiple sixth-point elements, thereby extending the mantissa of the sixth-point elements to 24 bits. The extension method can be to pad with zeros at the lower bits, specifically by appending multiple zero bits after the least significant bit of the mantissa, to reserve the precision margin required for subsequent rounding and normalization.
[0151] Meanwhile, the extended mantissa can participate in subsequent exponent alignment and right shift alignment, thus ensuring that the mantissa precision of the final fifth floating-point element meets the requirement of 23 valid mantissas in the FP32 format.
[0152] During the process of aligning the exponents in multiple sixth-point elements, the processing module can uniformly convert the extended exponents carried by different sixth-point elements into signed two's complement representations with the same bit width (e.g., 10 bits), and adapt their numerical range to the exponent comparison requirements of the FP32 accumulation stage.
[0153] Exponent alignment can be achieved by: selecting the largest exponent from the exponent values of multiple sixth-point elements as the alignment benchmark, and then performing a right shift operation on the mantissa of each sixth-point element based on the difference between the exponent of each sixth-point element and the largest exponent, while simultaneously generating the corresponding Guard, Round, and Sticky bits.
[0154] Specifically, the exponent alignment operation can be performed in conjunction with the mantissa expansion operation. Their connection is achieved through timing synchronization via shared control signals, which in turn provides a directly comparable exponent input to the exponent comparison module in the accumulation module 130. At the same time, the aligned exponent can participate in subsequent maximum exponent filtering and mantissa right shift alignment, which can ensure that multiple sixth floating-point elements and third floating-point elements are comparable and operable in the exponent dimension.
[0155] Specifically, during the mantissa expansion and exponent alignment of the sixth floating-point element, the processing module can extract the exponent field and mantissa field from each sixth floating-point element; perform parallel comparisons on the exponent fields of all sixth floating-point elements to determine the maximum exponent value; then, based on the difference between the exponent of each sixth floating-point element and the maximum exponent, control the right shift of the corresponding mantissa by a number of bits, and record the shifted-out bits to form the GRS field during the right shift; simultaneously, expand the right-shifted mantissa to the target bit width, with expansion methods including padding with zeros at low bits, sign expansion at high bits, or concatenation of GRS bits; and package the expanded mantissa, aligned exponent, and sign bit into a fourth operand according to a preset format, and send it to the accumulation module 130 for subsequent accumulation calculations.
[0156] As an example, in the (8×4)×(4×8)→(8×8) matrix multiplication and addition scenario, 16 groups of sixth floating-point elements are generated each cycle (each group corresponds to an FP32 output position). The processing module performs mantissa expansion and exponent alignment on the four sixth floating-point elements in each group. For example, when the exponents of the four seventh floating-point elements in a group are 102, 105, 103, and 105 respectively, the maximum exponent of 105 is taken as the reference, and the mantissas corresponding to exponents 102 and 103 are right-shifted by 3 bits and 2 bits respectively, and the corresponding GRS bits are generated. After the mantissa is right-shifted, it is uniformly expanded to 24 bits, which together with the aligned exponents form a well-formatted fourth operand for the accumulation module to perform high-precision accumulation with the FP32 initial value.
[0157] In this application, a processing module is added to the matrix multiplication and addition module 120 to perform mantissa expansion and exponent alignment operations on the seventh floating-point element output by the multiplication module. This eliminates the format gap between the multiplication result and the accumulated input. The mantissa is expanded to 24 bits and the low bits are padded with zeros, which can provide sufficient precision margin for subsequent addition operations and avoid the loss of low-bit information due to insufficient bit width.
[0158] In some embodiments, this application also provides a processor, including the BP16 matrix multiply-accumulate operation circuit based on the RISC-V architecture provided in this application.
[0159] In this application, the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0160] In some embodiments, this application also provides an electronic device, including the BP16 matrix multiply-accumulate circuit based on the RISC-V architecture provided in this application or the processor provided in this application.
[0161] In this embodiment, the electronic device may be equipped with a processor, which may be equipped with a BP16 matrix multiplication and addition operation circuit based on the RISC-V architecture and connected to the system bus to provide computing and control capabilities to support the operation of the electronic device.
[0162] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A BP16 matrix multiplication and addition operation circuit based on RISC-V architecture, characterized in that, include: The input module is configured to input operands, the operands including a first operand, a second operand and a third operand, the first operand including a plurality of first floating-point elements in BP16 format, the second operand including a plurality of second floating-point elements in BP16 format, and the third operand including a plurality of third floating-point elements in a preset format. A matrix multiplication and addition module is connected to the input module and is configured to periodically perform row and column multiplication operations on multiple first floating-point elements in a preset number of rows in the first operand and the second floating-point elements in the second operand in the form of a matrix to obtain a fourth operand, wherein the fourth operand includes multiple fourth floating-point elements in a preset format. An accumulation operation module is connected to the matrix multiplication and addition operation module and the input module, and is configured to accumulate multiple fourth floating-point elements with the third floating-point elements to obtain a fifth operand, wherein the fifth operand includes multiple fifth floating-point elements in a preset format; The output module is connected to the accumulation operation module and is configured to output the target operand based on multiple fifth floating-point elements; A preprocessing module, wherein the input end of the preprocessing module is connected to the input module, and the output end of the preprocessing module is connected to the matrix multiplication and addition operation module; The preprocessing module is configured to preprocess the first operand, the second operand, and the third operand respectively to obtain the preprocessed first operand, the second operand, and the third operand. The exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element are represented using signed two's complement. When the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element reach the minimum negative value represented by signed two's complement, independent protection processing is performed on the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element.
2. The BP16 matrix multiplication and addition circuit based on RISC-V architecture according to claim 1, characterized in that, The preprocessing module includes a splitting module, a special data detection module, and a hidden bit recovery module; The splitting module is connected to the input module, the hidden bit recovery module, the matrix multiplication and addition module, and the accumulation module, respectively. The special data detection module is connected to the input module, the hidden bit recovery module, and the accumulation module, respectively. The hidden bit recovery module is connected to the matrix multiplication and addition module and the accumulation module, respectively. The splitting module is configured to split the first floating-point element, the second floating-point element, and the third floating-point element into bits respectively, so as to obtain the sign bit, exponent, and mantissa in the first floating-point element, the second floating-point element, and the third floating-point element. The special data detection module is configured to perform normalization detection on the mantissas of the first floating-point element, the second floating-point element, and the third floating-point element respectively, and obtain normalization detection results. The hidden bit recovery module is configured to recover the hidden bit of the mantissa in the first floating-point element, the second floating-point element, and the third floating-point element based on the normalization detection result.
3. The BP16 matrix multiplication and addition circuit based on RISC-V architecture according to claim 2, characterized in that, The preprocessing module also includes a leading zero detection module, a mantissa normalization module, and an exponent adjustment module. The leading zero detection module is connected to the matrix multiplication and addition module, the accumulation module, the hidden bit recovery module, and the mantissa normalization module, respectively. The mantissa normalization module is connected to the matrix multiplication and addition module and the accumulation module, respectively. The exponent adjustment module is connected to the splitting module, the leading zero detection module, and the matrix multiplication and addition module, respectively. The leading zero detection module is configured to perform leading zero detection on the denormalized mantissas in the first floating-point element, the second floating-point element, and the third floating-point element respectively, to obtain the first leading zero detection result; The mantissa normalization module is configured to normalize the unnormalized mantissas in the first floating-point element and the second floating-point element based on the first leading zero detection result. Both the exponent adjustment module and the mantissa normalization module are configured to process the exponent and the denormalized mantissa in the third floating-point element based on the first leading zero detection result, so as to obtain the preprocessed third operand.
4. The BP16 matrix multiplication and addition circuit based on RISC-V architecture according to claim 1, characterized in that, The accumulation operation module includes an exponent comparison module, an alignment shift module, and an addition operation module; The exponent comparison module is connected to the preprocessing module, the alignment shift module, and the output module, respectively, and the addition module is connected to the preprocessing module, the alignment shift module, and the output module, respectively. The exponent comparison module is configured to filter out the largest exponent from the exponents of the floating-point elements in the preprocessed third operand and the exponents of the multiple fourth floating-point elements to obtain the exponent of the fifth floating-point element. The alignment shift module is configured to right-shift and align the mantissas of multiple fourth floating-point elements based on the maximum exponent, thereby obtaining multiple right-shifted and aligned mantissas. The addition module is configured to perform an addition operation on the mantissa of the floating-point element in the preprocessed third operand and multiple right-shifted mantissas to obtain the mantissa of the fifth floating-point element.
5. The BP16 matrix multiplication and addition circuit based on RISC-V architecture according to claim 4, characterized in that, The output module is configured to perform leading zero detection on the mantissas of the plurality of fifth floating-point elements to obtain a second leading zero detection result, and to normalize the mantissas of the plurality of fifth floating-point elements based on the second leading zero detection result to generate the mantissas in the target operand. Based on the second leading zero detection result, the exponents of multiple fifth floating-point elements are adjusted to generate the exponent in the target operand; Output the target operand.
6. The BP16 matrix multiplication and addition circuit based on RISC-V architecture according to claim 1, characterized in that, The matrix multiplication and addition module includes a multiplexer and a multiplication module; The multiplexer is connected to the input module and the multiplication module, respectively, and the multiplication module is connected to the accumulation module. The multiplexer is configured to periodically output a plurality of the first floating-point elements in a preset number of rows of the first operand; The multiplication module is configured to perform row and column multiplication operations on multiple first floating-point elements in a preset number of rows in the first operand and second floating-point elements in the second operand in the form of a matrix, to obtain multiple fourth floating-point elements.
7. The BP16 matrix multiplication and addition circuit based on RISC-V architecture according to claim 6, characterized in that, The matrix multiplication and addition module also includes a processing module; The processing module is connected to the accumulation operation module; The multiplication module is configured to perform row and column multiplication on multiple first floating-point elements in a preset row number in the first operand and second floating-point elements in the second operand in the form of a matrix to obtain multiple sixth floating-point elements. The processing module is configured to expand the mantissas of the plurality of sixth floating-point elements and align the exponents of the plurality of sixth floating-point elements to obtain the fourth operand.
Citation Information
Patent Citations
Apparatus, method and system for 8-bit floating point matrix dot product instructions
CN118605946A
Multi-precision matrix calculation unit and use method thereof
CN121167096A
Floating point multiply-accumulate unit facilitating variable data precision
CN121420281A