Matrix multiplication and addition operation circuit based on risc-v architecture
Patent Information
- Application Number
- CN202610508871.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-17
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-04-17
AI Technical Summary
[0005]针对现有技术的不足,本申请提供了一种基于RISC-V架构的矩阵乘加运算电路,解决了不同低精度浮点格式难以在同一硬件架构下统一支持的技术问题,显著提升了硬件复用率、能效比与指令兼容性,并可以与RISC-V向量扩展高度融合,适用于嵌入式AI加速器及通用处理器核
[0018]本申请提供的基于RISC-V架构的矩阵乘加运算电路,包括输入模块、矩阵乘加运算模块、累加运算模块及输出模块,通过输入模块对操作数浮点元素执行精度识别,实现了FP8与BP16格式的自动判别,矩阵乘加运算模块基于识别信息执行行列乘法运算,强制约束第一浮点元素与第二浮点元素同属单一精度格式,保障数值范围匹配与运算一致性;累加运算模块在统一预设格式下完成第四浮点元素与第三浮点元素的逐元素累加,规避跨格式转换带来的控制复杂性与延迟开销;输出模块完成尾数规格化与指数调整,生成符合目标格式的目标操作数,整个电路结构紧凑、控制逻辑清晰,可无缝嵌入RISC-V向量处理器核,有效支撑AI推理与训练负载,显著降低硬件冗余与系统设计复杂度。
Smart Images

Figure CN122044517B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of processor microarchitecture and artificial intelligence computing acceleration technology, and in particular to a matrix multiplication and addition circuit based on the RISC-V architecture. Background Technology
[0002] With the rapid development of artificial intelligence and high-performance computing technologies, matrix multiplication and addition operations have become the core computing load in modern processors and accelerators, and are widely used in scenarios such as deep learning training and inference, scientific computing and big data analysis.
[0003] To balance computational density and energy efficiency, low-precision floating-point formats such as BF16 (Brain Precision 16) and FP8 (Floating-Point 8) are gradually replacing traditional FP32 / FP16 as the mainstream data representation method for AI workloads. BF16 ensures a wide dynamic range by retaining the 8-bit exponent of FP32, while FP8 is further compressed to 8 bits (such as E4M3 / E5M2), significantly improving bandwidth utilization and computational throughput.
[0004] However, current mainstream implementations mostly use dedicated hardware paths to support BF16 and FP8 respectively, resulting in independent data paths, control logic, and special value processing mechanisms. Hardware resources are difficult to reuse, system integration is complex, and scalability is limited. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this application provides a matrix multiplication and addition circuit based on the RISC-V architecture, which solves the technical problem that different low-precision floating-point formats are difficult to support uniformly under the same hardware architecture. It significantly improves hardware reuse rate, energy efficiency and instruction compatibility, and can be highly integrated with RISC-V vector extensions, making it suitable for embedded AI accelerators and general-purpose processor cores.
[0006] In a first aspect, this application provides a matrix multiplication and addition circuit based on a RISC-V architecture, comprising: The input module is configured to input operands and perform precision recognition on the floating-point elements in the input operands to obtain recognition information; the operands include a first operand, a second operand, and a third operand, and the third operand includes multiple third floating-point elements in a preset format; The matrix multiplication and addition module is connected to the input module and is configured to periodically perform matrix multiplication on multiple first floating-point elements in a preset number of rows of the first operand and second floating-point elements in the second operand, based on recognition information, to obtain a fourth operand. The fourth operand includes multiple fourth floating-point elements in a preset format. The first and second floating-point elements are either floating-point numbers in FP8 format or floating-point numbers in BP16 format. The accumulation operation module connects the matrix multiplication and addition operation module and the input module, and is configured to accumulate multiple fourth floating-point elements with third floating-point elements to obtain a fifth operand, which includes multiple fifth floating-point elements in a preset format. The output module is connected to the accumulation module and is configured to output the target operand based on multiple fifth floating-point elements.
[0007] In one embodiment, the matrix multiplication and addition circuit based on the RISC-V architecture further includes a preprocessing module; The input end of the preprocessing module is connected to the input module, and the output end of the preprocessing module is connected to the matrix multiplication and addition operation module. The preprocessing module is configured to preprocess the first operand, the second operand, and the third operand respectively to obtain the preprocessed first operand, the second operand, and the third operand.
[0008] In one embodiment, the exponents of the floating-point elements in the third and fourth floating-point elements after preprocessing are represented using signed two's complement.
[0009] In one embodiment, when the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element reach the minimum negative value represented by the two's complement, independent protection processing is performed on the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element.
[0010] In one embodiment, the preprocessing module includes a splitting module, a special data detection module, and a hidden bit recovery module; The splitting module is connected to the input module, the hidden bit recovery module, the matrix multiplication and addition module, and the accumulation module, respectively. The special data detection module is connected to the input module, the hidden bit recovery module, and the accumulation module, respectively. The hidden bit recovery module is connected to the matrix multiplication and addition module and the accumulation module, respectively. The splitting module is configured to split the first floating-point element, the second floating-point element, and the third floating-point element into bits respectively, so as to obtain the sign bit, exponent, and mantissa in the first floating-point element, the second floating-point element, and the third floating-point element. The special data detection module is configured to perform normalization detection on the mantissas of the first floating-point element, the second floating-point element, and the third floating-point element respectively, and obtain the normalization detection results. The hidden bit recovery module is configured to recover the hidden bit of the mantissa in the first, second, and third floating-point elements based on the normalization detection results.
[0011] In one embodiment, the preprocessing module further includes a leading zero detection module, a mantissa normalization module, and an exponent adjustment module; Among them, the leading zero detection module is connected to the matrix multiplication and addition operation module, the accumulation operation module, the hidden bit recovery module and the mantissa normalization module respectively; the mantissa normalization module is connected to the matrix multiplication and addition operation module and the accumulation operation module respectively; and the exponent adjustment module is connected to the splitting module, the leading zero detection module and the matrix multiplication and addition operation module respectively. The leading zero detection module is configured to perform leading zero detection on the denormalized mantissas in the first floating-point element, the second floating-point element, and the third floating-point element respectively, and obtain the first leading zero detection result; The mantissa normalization module is configured to normalize the unnormalized mantissas in the first floating-point element and the second floating-point element based on the first leading zero detection result. Both the exponent adjustment module and the mantissa normalization module are configured to process the exponent and the denormalized mantissa in the third floating-point element based on the first leading zero detection result, so as to obtain the preprocessed third operand.
[0012] In one embodiment, the accumulation module includes an exponent comparison module, an alignment shift module, and an addition module; The exponent comparison module is connected to the preprocessing module, the alignment and shifting module, and the output module, respectively; the addition module is connected to the preprocessing module, the alignment and shifting module, and the output module, respectively. The exponent comparison module is configured to filter the largest exponent from the exponents of the floating-point elements in the preprocessed third operand and the exponents of multiple fourth floating-point elements to obtain the exponent of the fifth floating-point element. The alignment shift module is configured to right-shift and align the mantissas of multiple fourth floating-point elements based on the maximum exponent, resulting in multiple right-shifted and aligned mantissas. The addition module is configured to add the mantissa of the floating-point element in the preprocessed third operand to multiple right-shifted mantissas to obtain the mantissa of the fifth floating-point element.
[0013] In one embodiment, the output module is configured to perform leading zero detection on the mantissas of a plurality of fifth floating-point elements to obtain a second leading zero detection result, and to normalize the mantissas of the plurality of fifth floating-point elements based on the second leading zero detection result to generate the mantissa in the target operand; to adjust the exponents of the plurality of fifth floating-point elements based on the second leading zero detection result to generate the exponent in the target operand; and to output the target operand.
[0014] In one embodiment, the matrix multiply-add operation module includes a multiplexer and a multiplication operation module; The multiplexer is connected to the input module and the multiplication module, respectively, and the multiplication module is connected to the accumulation module. The multiplexer is configured to periodically output multiple first floating-point elements of a preset number of rows in the first operand based on identification information; The multiplication module is configured to perform matrix multiplication on multiple first floating-point elements in a preset row number in the first operand and second floating-point elements in the second operand to obtain multiple fourth floating-point elements.
[0015] In one embodiment, the matrix multiplication and addition module further includes a processing module; The processing module is connected to the accumulation operation module; The multiplication module is configured to perform row and column multiplication on multiple first floating-point elements in a preset row number in the first operand and second floating-point elements in the second operand in the form of a matrix, to obtain multiple sixth floating-point elements; The processing module is configured to expand the mantissas of multiple sixth-point elements and align the exponents of multiple sixth-point elements to obtain a fourth operand.
[0016] Secondly, this application also provides a processor, including the matrix multiplication and addition circuit based on the RISC-V architecture provided in the first aspect.
[0017] Thirdly, this application also provides an electronic device, including the matrix multiplication and addition circuit based on the RISC-V architecture provided in the first aspect or the processor provided in the second aspect.
[0018] The matrix multiplication and addition circuit based on the RISC-V architecture provided in this application includes an input module, a matrix multiplication and addition module, an accumulation module, and an output module. The input module performs precision identification on the floating-point elements of the operands, realizing automatic discrimination between FP8 and BP16 formats. The matrix multiplication and addition module performs row and column multiplication operations based on the identification information, forcibly constraining the first and second floating-point elements to belong to the same single-precision format, ensuring numerical range matching and operational consistency. The accumulation module completes the element-wise accumulation of the fourth and third floating-point elements under a unified preset format, avoiding the control complexity and latency overhead caused by cross-format conversion. The output module completes mantissa normalization and exponent adjustment to generate target operands that conform to the target format. The entire circuit has a compact structure and clear control logic, and can be seamlessly embedded into the RISC-V vector processor core, effectively supporting AI inference and training workloads, and significantly reducing hardware redundancy and system design complexity. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 A first schematic block diagram of a matrix multiplication and addition circuit based on the RISC-V architecture provided in an embodiment of this application; Figure 2 A second schematic block diagram of a matrix multiplication and addition circuit based on the RISC-V architecture provided in an embodiment of this application; Figure 3 A first architecture diagram of a matrix multiplication and addition circuit based on RISC-V architecture provided for embodiments of this application; Figure 4 A second architecture diagram of a matrix multiplication and addition operation circuit based on RISC-V architecture provided for embodiments of this application; Figure 5 A third architecture diagram of a matrix multiplication and addition circuit based on RISC-V architecture provided for embodiments of this application; Figure 6 The fourth architecture diagram of the matrix multiplication and addition circuit based on the RISC-V architecture provided in the embodiments of this application is shown. Detailed Implementation
[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0022] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0023] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0024] It should also be further understood that the term “and / or” as used in this application specification and the appended claims is to mean any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0025] Furthermore, in this application, unless otherwise explicitly specified or limited in the embodiments, the terms "installation," "connection," "joining," and "fixing" appearing in the embodiments should be interpreted broadly. For example, a connection can be a fixed connection, a detachable connection, or an integral part; it can also be a mechanical connection, an electrical connection, etc. Of course, it can also be a direct connection, or an indirect connection through an intermediate medium, or it can be the internal communication between two components, or the interaction between two components. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific implementation.
[0026] In related technologies, deep learning models are continuously increasing their demands for energy efficiency and throughput in low-bit floating-point operations in scenarios such as artificial intelligence inference and training, and high-performance vector computing. Currently, FP32 or FP16 are mostly used for matrix multiplication and addition operations. Although they have high numerical accuracy, they have significant bottlenecks in terms of computing power density, power consumption control, and storage bandwidth. In particular, under the RISC-V architecture, the general floating-point pipeline is difficult to efficiently support large-scale parallel matrix operations in FP8 and BP16 formats. Furthermore, the handling of denormalized numbers (Denorm), NaN, Inf, and exponential boundary anomalies is scattered across multiple logic units, resulting in high hardware complexity, high latency, and low energy efficiency.
[0027] Furthermore, during the accumulation process from FP8×FP8 or BP16×BP16 to FP32, due to the large differences in the dynamic range of the exponent, the narrowness of the mantissa, and the tight coupling between rounding and normalization, precision loss or overflow errors are prone to occur, and there is a lack of a unified and predictable hardware co-processing mechanism.
[0028] To address this, this application provides a matrix multiplication and addition circuit based on the RISC-V architecture, which achieves unified hardware support for two low-precision floating-point formats, FP8 and BP16, significantly improving computing power density and energy efficiency in AI scenarios. It natively supports dynamic recognition of dual formats and data path reuse, reducing data migration and bandwidth pressure. It is highly integrated with the RISC-V vector architecture, providing unified and predictable hardware processing for special values and exponential boundaries, and is suitable for CPUs, AI accelerators, and heterogeneous computing platforms.
[0029] In the matrix multiplication and addition circuit based on the RISC-V architecture, three operands are received through the input module, including the first operand, the second operand, and the third operand, which enables efficient processing of low-precision floating-point data.
[0030] This application solves the problem of lack of dedicated hardware support for low-precision matrix multiplication and addition by periodically performing matrix multiplication on multiple first floating-point elements in a preset number of rows in the first operand and second floating-point elements in the second operand to generate intermediate results. This significantly improves computational efficiency.
[0031] The accumulation module adds the intermediate result to the third operand, realizing the precision conversion from low-precision multiplication to FP32, effectively avoiding the accumulation of numerical errors caused by insufficient precision. The output module outputs the target operand based on the accumulation result, ensuring the integrity and usability of the calculation result.
[0032] Furthermore, this application also uses a preprocessing module to perform format adaptation and pre-normalization on the input operands, solving the problem of handling special values such as denormalized numbers, NaN, and Inf in FP8 and BP16 data, and ensuring the numerical consistency of the calculations. At the same time, the exponents of the floating-point elements in the third and fourth floating-point elements of the preprocessed operands are represented using signed two's complement, simplifying the exponent comparison logic, supporting wide dynamic range mapping, and ensuring correct alignment even in the smallest exponent scenario.
[0033] Meanwhile, independent protection is performed on the minimum negative exponent, eliminating the risk of numerical collapse under the worst boundary and ensuring the determinism and verifiability of the calculation results in the entire input domain.
[0034] The collaborative work of the splitting module, special data detection module, and hidden bit recovery module enables unified parsing of both FP8 and BP16 formats, explicitly exposing hidden bits and avoiding loss of mantissa precision.
[0035] The coordinated use of the leading zero detection module, the mantissa normalization module, and the exponent adjustment module eliminates computational redundancy caused by denormalized numbers, unifies the exponent reference system, and ensures the accuracy of accumulation alignment.
[0036] The combination of the exponent comparison module, alignment shift module, and addition operation module ensures that the numerical accuracy of the accumulation process is not degraded, supports multi-channel parallel accumulation, and is compatible with the GRS bit generation required by the RISC-V rounding mode.
[0037] Meanwhile, the output module can also perform leading zero detection, normalization shift, and dynamic exponent adjustment to ensure that the output strictly conforms to the IEEE 754FP32 format, supports five rounding modes of RISC-V, and realizes automatic setting of the exception flag.
[0038] The multiplexer, multiplication module, and floating-point number work together to achieve high-throughput matrix multiplication, support RISC-V vector instruction periodic issuance, and reduce the temporary storage area of multiplication results.
[0039] The processing module performs mantissa expansion and exponent alignment on the floating-point elements output by the multiplication module, eliminating the data format gap between the multiplication and accumulation modules and avoiding truncation errors caused by bit width mismatch.
[0040] This application can significantly improve computing power density and energy efficiency in AI scenarios, natively support dual formats, reduce data migration and bandwidth pressure, and is highly integrated with the RISC-V vector architecture, providing unified and predictable hardware processing for special values and exponential boundaries. It is suitable for CPUs, AI accelerators and heterogeneous computing platforms.
[0041] Please see Figure 1 , Figure 1 This is a first schematic block diagram of a matrix multiplication and addition circuit based on the RISC-V architecture provided in an embodiment of this application. Figure 1 As shown, this application provides a matrix multiplication and addition circuit based on the RISC-V architecture, including: The input module 110 is configured to input operands and perform precision recognition on the floating-point elements in the input operands to obtain recognition information; the operands include a first operand, a second operand, and a third operand, and the third operand includes multiple third floating-point elements in a preset format; The matrix multiplication and addition module 120 is connected to the input module 110 and is configured to periodically perform row and column multiplication and addition operations on multiple first floating-point elements of a preset number of rows in the first operand and second floating-point elements of the second operand in the form of a matrix, based on recognition information, to obtain a fourth operand; the fourth operand includes multiple fourth floating-point elements of a preset format; the first floating-point elements and the second floating-point elements are both floating-point numbers in FP8 format, or the first floating-point elements and the second floating-point elements are both floating-point numbers in BP16 format; The accumulation operation module 130 is connected to the matrix multiplication and addition operation module 120 and the input module 110, and is configured to accumulate the fourth floating-point element and the third floating-point element to obtain the fifth operand, which includes multiple fifth floating-point elements in a preset format. Output module 140 is connected to accumulation module 130 and is configured to output target operands based on multiple fifth floating-point elements.
[0042] Specifically, this application provides a hardware native adaptation approach for the RISC-V Vector Extension (RVV) instruction set. Specifically, it can use FP32+=FP8×FP8 or FP32+=BP16×BP16 as the basic computation paradigm, modularly decouple and periodically pipeline the input data stream, matrix operation structure, accumulation alignment logic and output regularization process, and achieve high throughput, low power consumption and strong robust low-precision matrix multiplication and addition hardware acceleration while ensuring IEEE754 compatibility and RISC-V rounding semantic consistency.
[0043] The input module 110 can be understood as an interface circuit for receiving three 512-bit wide data streams from the RISC-V vector register group (such as v0–v31), and its input port supports data transfer triggered by RISC-V vector loading instructions (such as vlw.v).
[0044] Both the first and second operands are in low-precision floating-point format, and include either FP8 format (E4M3 or E5M2) or BP16 format (1 sign bit + 8 exponent + 7 mantissa, with an exponent bias of 127). The two formats can be dynamically switched using a control signal. The first and second operands can be understood as... Figure 4 and Figure 6 The opa and opb in the text.
[0045] The third operand is in FP32 format, which corresponds to the RISC-V standard single-precision floating-point format (IEEE 754), including 1 sign bit, 8 exponent bits, and 23 mantissa bits.
[0046] Specifically, the input module 110 contains a buffer register array to temporarily store the block data of each operand in order to match the periodic processing rhythm of subsequent modules. The buffer depth can be set to 2–4 cycles according to the actual timing constraints, such as a double buffer structure or a four-cycle FIFO.
[0047] The input module 110 is configured to perform precision identification on the floating-point elements in the input operands to obtain identification information. The input module 110 can integrate a precision identification unit, which can be a combinational logic circuit based on a control signal decoder or instruction field parser. It is configured to identify the low-precision floating-point format used in the current operation based on the value of the vtype field specified in the RISC-V vector instruction encoding or a dedicated format configuration register (such as fmatmul_fmtCSR).
[0048] When the identification information indicates FP8 format, the input module 110 outputs the corresponding bit width parameters (1 sign bit, 4 or 5 exponent bits, 3 or 2 mantissa bits), exponent bias value (7 or 15), and hidden bit presence flag for FP8; when the identification information indicates BP16 format, the input module 110 outputs the corresponding bit width parameters (1 sign bit, 8 exponent bits, 7 mantissa bits), exponent bias value 127, and hidden bit presence flag for BP16.
[0049] In this application, the identification information can be used as a global configuration signal and synchronously distributed to subsequent modules to dynamically select the data path width, exponent alignment logic and rounding control mode, so that only the format discrimination and parameter derivation functions can be performed without changing the input data itself.
[0050] Meanwhile, the recognition process can be completed in a single cycle, and it supports maintaining the same recognition result in multiple consecutive matrix multiplication and addition cycles to improve throughput efficiency; the input module 110 is compatible with the RISC-V Vector Extension (RVV) instruction set and can trigger recognition actions in response to vfmatmul.vv class instructions.
[0051] The input module 110 is configured to input operands and performs precision recognition on the floating-point elements in the input operands to obtain recognition information. This information can then be used to construct a format-aware starting point, enabling all subsequent operation modules to be parameterized based on the same recognition result.
[0052] The input module 110 forms a closed-loop driving relationship with the RISC-V Vector Extension (RVV) instruction set, which can complete format discrimination and path configuration without software intervention. Thus, without sacrificing numerical robustness, it can achieve seamless switching and high resource reuse between BF16 and FP8 under the same hardware framework.
[0053] The matrix multiplication and addition module 120 can be understood as an operation unit that performs low-precision floating-point matrix multiplication. It is configured as the core calculation unit to perform FP8×FP8 or BP16×BP16 to generate intermediate FP32 format results, and completes the calculation of all 64 output elements through a 4-cycle time-division multiplexing structure.
[0054] Among them, such as Figure 3 and Figure 5 As shown, each cycle processes 2 rows of results, each row containing 8 FP8×FP8 or 4 BP16×BP16 multiplication results; the matrix multiplication and addition module 120 may include a multiplexer and a multiplication module. The multiplexer periodically selects 16 FP8 elements or 8 BP16 elements from a preset number of rows (e.g., 2 rows) from the first operand, and pairs them with the 8 FP8 elements or 4 BP16 elements in the corresponding column of the second operand.
[0055] The multiplication module performs floating-point multiplication on each pair of FP8 or BP16 elements and outputs an intermediate result with an extended bit width. The sign bit is obtained by XORing the two input sign bits, the exponent is obtained by subtracting the corresponding bias from the sum of the two input exponents and correcting it with leading zeros, and the mantissa is obtained by multiplying the two input mantissas (including the hidden bit after recovery) and extended to at least 16 bits.
[0056] The extended result format can be Sgn(1) + Exp(10) + Man(24), where Exp(10) is represented by signed two's complement to cover the wide range of exponents that may occur after multiplication of FP8 and BP16.
[0057] The matrix multiplication and addition module 120 is configured to periodically perform matrix multiplication on multiple first floating-point elements of a preset number of rows in the first operand and second floating-point elements in the second operand, based on identification information, to obtain a fourth operand. Its function is to build a unified intermediate representation hub. By forcing the first and second floating-point elements to be isomorphic (both FP8 or both BP16) during the operation phase, it avoids the risk of exponent mismatch and mantissa overflow caused by heterogeneous multiplication. The matrix multiplication and addition module 120 strictly follows the configuration path driven by identification information and does not introduce cross-format mixed operations, which can ensure the numerical stability during the multiplication phase.
[0058] The accumulation operation module 130 can be understood as a dedicated circuit that completes the alignment, summation and anomaly detection of the intermediate result of FP32 with the initial accumulated value (third operand); its input terminals receive the fourth operand output from the matrix multiplication and addition operation module 120 and the third operand output from the input module 110, respectively.
[0059] The accumulation operation module 130 may include an exponent comparison module, an alignment shift module, and an addition operation module; the exponent comparison module is configured to compare the exponent of a fourth floating-point element corresponding to each output position with the exponent of a third floating-point element, and select the largest exponent between the two as the exponent benchmark of the fifth floating-point element at that position.
[0060] The alignment shift module is configured to right-shift the mantissa of the operand corresponding to the smaller exponent based on the largest exponent. The number of right shifts is equal to the exponent difference, and guard, round, and sticky bits are generated in the shifted-out low bits, i.e., G bits / R bits / S bits.
[0061] The addition module is configured to perform signed addition on two aligned mantissas (including sign extension) and output the mantissa of the fifth floating-point element. The addition module supports a NaN propagation mechanism: if any input is DQNaN, it directly outputs DQNaN without performing subsequent alignment and addition; its hardware implementation uses a dual-path structure, where the normal path performs aligned addition, and the abnormal path directly leads to NaN / Inf, to reduce critical path latency.
[0062] The accumulation module 130 is configured to accumulate multiple fourth floating-point elements with third floating-point elements to obtain a fifth operand. The accumulation module 130 eliminates the complex conversion logic required for cross-format alignment by accumulating multiple fourth floating-point elements with third floating-point elements under the same preset format.
[0063] Meanwhile, the accumulation module 130 requires the third floating-point element and the fourth floating-point element to have the same preset format. Therefore, their exponent bit width, mantissa bit width and implicit bit rules are completely consistent, so that alignment and addition can be completed in the same data path without the need for format conversion logic.
[0064] In addition, each fifth floating-point element in the fifth operand output by the accumulation module 130 retains the preset format, with its exponent equal to the maximum exponent and its mantissa being the result of addition, which can provide an input basis for the normalization processing of the subsequent output module 140.
[0065] The output module 140 can be understood as a final processing unit that performs normalization, rounding, overflow judgment and format encapsulation on the fifth floating-point element; the input of the output module 140 can receive 64 FP32 elements of the fifth operand; at the same time, the output module 140 includes a leading zero detection module, a normalization shift module, a rounding module and an exception flag generation module.
[0066] The leading zero detection module is configured to perform a leading zero count (LZC) on the mantissa of each fifth floating-point element to determine the number of bits to be left-shifted; the normalization shift module is configured to perform a left shift on the mantissa based on the LZC result, so that the most significant bit is at the 23rd bit (i.e., in the form of 1.xxxx), and at the same time perform a corresponding subtraction adjustment on the exponent based on the number of bits to be left-shifted.
[0067] The rounding module is configured to generate a rounding increment based on the GRS bit and the current RISC-V rounding mode (RNE / RTZ / RDN / RUP / RMM), and add it to the normalized mantissa. If a carry occurs, a second right rounding is triggered (Exp+1, Man shifts right by 1 bit).
[0068] The exception flag generation module is configured to output five types of exception flags: NV, DZ, OF, UF, and NX, based on the final exponent and mantissa combination. The target operands output by the exception flag generation module are in standard FP32 format, which can be directly written back to the RISC-V vector register file and supports write-back using memory instructions such as vsw.v.
[0069] The output module 140 is configured to output target operands based on multiple fifth-point elements, enabling final normalization and FP32 mapping, achieving seamless integration between low-precision input and high-precision output. The target operands output by this module are in FP32 format, with the sign bit directly inherited from the fifth-point element. The exponent is compensated and mapped to FP32 bias 127, and the mantissa is rounded and truncated or extended to 23 decimal places. Simultaneously, the output module 140 supports the RISC-V floating-point exception handling mechanism, allowing exception flags to be written to the fcsr register for software lookup.
[0070] In this embodiment, when the RISC-V processor executes the vfmatmul.vv instruction, the input module 110 loads the first operand (vs1), the second operand (vs2), and the third operand (vs3) in parallel from the vector register. All three are 512 bits wide. The matrix multiply-add module 120 is scheduled by cycle. In the first cycle, it reads rows 0-1 of vs1 (16 FP8s or 8 BP16s) and all 8 columns of vs2 (64 FP8s or 32 BP16s). After multiplexing and multiplication array processing, it generates 128... The intermediate result of one FP8×FP8 or 32 BP16×BP16 product (i.e., rows 0–1 of the fourth operand); the accumulation module 130 synchronously receives this intermediate result and the FP32 initial value at the corresponding position of vs3, performs exponent comparison, right shift alignment, and mantissa addition, generating 16 fifth floating-point elements; the output module 140 performs LZC detection, left normalization, rounding, and overflow judgment on it, generating 16 target FP32 results and writing them to the destination register; subsequent cycles process the remaining rows sequentially, completing all 64 outputs in 4 cycles. Throughout the entire process, automatic FP8 / BP16 format adaptation, Denorm front normalization, NaN pass-through, and minimum exponent protection are all completed autonomously by hardware without software intervention.
[0071] As an example, taking the convolution kernel weight matrix ([8×8]) and activation feature matrix ([8×8]) of a certain layer in the ResNet-50 model as an example, both are quantized in FP8-E4M3 format and stored in the vector register; the third operand is the result of the previous accumulation of this layer, stored in FP32 format; after the circuit starts, the input module 110 loads vs1[0:1] (weights of the 0th-1st row), vs2 (all 8 columns of features) and vs3[0:1] (weights of the 0th-1st row) in cycle0. –1 row of initial accumulation); matrix multiplication and addition module 120 calculates the product between floating-point elements; accumulation module 130 adds the 8 products to the corresponding elements of vs3[0:1] respectively; output module 140 completes normalization and rounding, and outputs 16 FP32 results to vd[0:1]; cycle1–3 completes the calculation of the remaining rows in the same way, and finally obtains the complete [8×8] FP32 output matrix in vd. The whole process only requires 4 vector instruction cycles and no additional scalar intervention.
[0072] In this application, the input module 110 supports dynamic switching between FP8 and BP16 dual formats, thereby being compatible with the different requirements of different AI models for numerical accuracy and dynamic range, thus improving hardware versatility; the matrix multiplication and addition module 120 adopts a 4-cycle time-sharing multiplexing structure and integrates a parallel reduction tree, which can control the total delay of (8×8)×(8×8) matrix multiplication and addition within 4 cycles while maintaining a 512-bit vector bandwidth, significantly improving the computing power density per unit area.
[0073] Meanwhile, the accumulation module 130 has built-in exponent comparison and alignment shift hardware, and the output module 140 integrates LZC and rounding linkage logic, which avoids the overhead of multiple memory accesses or calls to scalar units for exponent adjustment in traditional solutions, and reduces overall power consumption.
[0074] In addition, each module sets up independent hardware paths for NaN, Inf, Denorm, and the minimum exponent, thus achieving deterministic exception handling and predictable timing, and meeting the strict consistency requirements of RISC-V vector extension for floating-point semantics. At the same time, all modules are designed based on the RISC-V vector register interface, which can be seamlessly integrated into the RISC-V CPU core or AI coprocessor, providing underlying hardware support for the RVV instruction set.
[0075] In some embodiments, such as Figure 2 As shown, the matrix multiplication and addition circuit based on the RISC-V architecture also includes a preprocessing module 150; wherein, the input terminal of the preprocessing module 150 is connected to the input module 110, and the output terminal of the preprocessing module 150 is connected to the matrix multiplication and addition module 120; the preprocessing module 150 is configured to preprocess the first operand, the second operand, and the third operand respectively to obtain the preprocessed first operand, the second operand, and the third operand.
[0076] In this embodiment, the preprocessing module 150 can be understood as a hardware logic unit that performs unified format adaptation, numerical normalization and anomaly prediction on the original input operands before the low-precision matrix multiplication and addition operation is started.
[0077] The preprocessing module 150 can convert raw data that cannot be directly used for high-precision accumulation, such as heterogeneous formats (FP8's E4M3 / E5M2 and BP16), denormalized numbers (Denorm), missing hidden bits, and special values (NaN / Inf / Zero), into an intermediate representation that meets the stable operation requirements of the subsequent matrix multiplication and accumulation modules.
[0078] Meanwhile, the preprocessing module 150 and the input module 110 form a serial data path, and its output strictly serves the pre-constraints proposed by the matrix multiplication and addition module 120 on operand format, exponent alignment and mantissa integrity.
[0079] The first operand is a multiplicand in FP8 or BP16 format. Before entering the preprocessing module 150, it has not undergone hidden bit recovery, exponent offset alignment, or normalization of denormalized numbers. The preprocessing module 150 can perform bit splitting, normalization detection, hidden bit recovery, and optional exponent adjustment on the first operand, and then output the preprocessed first operand. In this output data, the mantissa has been padded with hidden bits, the exponent has been mapped to a uniform extended bit width (not less than 10 bits), and the denormalized number has been left-normalized and the exponent has been corrected synchronously, thereby ensuring that it can perform lossless row and column multiplication operations with other operands in the matrix multiplication and addition operation module 120.
[0080] The second operand is a multiplier in FP8 or BP16 format, with the same structure as the first operand, also using E4M3 / E5M2 or BP16 format. The preprocessing module 150 performs a preprocessing process on it that is completely symmetrical to that of the first operand, including sign / exponent / mantissa separation, denorm detection and normalization, hidden bit insertion, and exponent field expansion. This processing ensures that the mantissa precision and exponent dynamic range of the second operand are consistent with those of the first operand. The two can achieve element-wise alignment and matching during the matrix multiplication stage, avoiding product truncation or overflow caused by asymmetry in mantissa bit width or difference in exponent bias.
[0081] The third operand is an initial value in FP32 format, whose original exponent and mantissa bit widths are significantly higher than those of FP8 or BP16 operands. The preprocessing module 150 performs exponent field remapping and mantissa zero extension on the third operand to keep its exponent representation compatible with the intermediate results generated by the first and second operands after preprocessing (such as Sgn(1) + Exp(10) + Man(24)).
[0082] Specifically, the preprocessing module 150 can map the 8-bit exponent of FP32 to a 10-bit signed two's complement representation through bias conversion, and extend the 23-bit mantissa by padding the high bits with zeros to at least 16 bits. This ensures that it can achieve bit-level alignment with the fourth operand (i.e., the intermediate result of FP32 output by matrix multiplication) in the subsequent accumulation operation module 130 in the exponent comparison, right shift alignment and mantissa addition reduction stages, preventing the loss of low-bit information or rounding deviation caused by bit width mismatch.
[0083] The preprocessing module 150 operates as follows: When the input module 110 sends the first operand, the second operand, and the third operand into the preprocessing module 150 in parallel, the module first obtains the sign bit, original exponent, and original mantissa of each operand through bit splitting logic. Subsequently, the special data detection module performs NaN / Inf / Zero / Denorm identification on the three data streams respectively. For data identified as Denorm, the leading zero detection module outputs the number of leading zeros, and the mantissa normalization module performs a left shift on the mantissa accordingly, while the exponent adjustment module synchronously corrects the exponent. For normalized data, the hidden bit recovery module inserts a hidden bit 1 before the most significant bit of the mantissa. Finally, all three operands are converted into an intermediate data stream with a unified exponent representation width (10-bit signed two's complement), an extended mantissa bit width (≥16 bits), and a standardized sign processing mechanism, and are synchronously output to the corresponding input port of the matrix multiplication and addition module 120 according to the periodic beat.
[0084] As an example, when the RISC-V vector instruction triggers the operation fp32+=fp8×fp8, vector register v0 provides 64 bytes of FP8 multiplicand (8×8 E4M3), v1 provides 64 bytes of FP8 multiplier (8×8 E5M2), and v2–v5 provide a total of 256 bytes of FP32 accumulation initial value (4×64 bytes). After receiving the data from v0 / v1 / v2, the preprocessing module 150 completes the parallel splitting and denorm detection of the three data streams in the first cycle; the second cycle starts the hidden bit recovery and mantissa normalization; the third cycle completes the exponent remapping and mantissa expansion; and the fourth cycle outputs the three preprocessed data streams to the matrix multiplication and addition module 120 in the form of a 512-bit wide bus. At this time, the mantissas of the first and second operands are both 16 bits and the exponents are both 10-bit signed two's complements, while the mantissa of the third operand is 24 bits and the exponent is 10-bit signed two's complements, all of which meet the consistency requirements of the matrix multiplication and addition and accumulation modules for the input data format.
[0085] In this application, by adding a preprocessing module 150 and performing format unification and numerical regularization on the three operands, the precision of the mantissa and the compatibility of the exponent dynamic range are ensured between the first operand and the second operand during FP8×FP8 or BP16×BP16 multiplication. At the same time, the preprocessing module 150 performs exponent remapping and mantissa expansion on the third operand, which enables the initial value of FP32 accumulation to be aligned with the intermediate result of FP8 or BP16 multiplication output without error in the subsequent accumulation stage.
[0086] In addition, the preprocessing module 150 can have built-in special data detection and independent protection logic, which can intercept Denorm, NaN and extremely small exponent scenarios in advance, and avoid them from causing chain anomalies in the main matrix multiplication and addition path, thereby improving the numerical robustness and operational predictability of the entire low-precision matrix multiplication and addition operation circuit.
[0087] In some embodiments, the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element are represented using signed two's complement.
[0088] In this embodiment, the exponent of the floating-point element in the third operand after preprocessing can be understood as the exponent value obtained after format conversion, normalization and bias adjustment by the preprocessing module 150. The exponent value is represented in signed two's complement form with a bit width of not less than 10 bits. Its function is to provide the subsequent exponent alignment module with a data format that can directly participate in arithmetic comparison and difference operation, avoiding the maximum value selection error caused by the distortion of negative exponent mapping under unsigned encoding. At the same time, the exponent and the exponent of the fourth floating-point element maintain the same format in the data path, and the two are jointly input to the exponent comparison module, forming a unified input basis for the exponent alignment operation.
[0089] In the third operand after preprocessing, the exponent of the floating-point element is an integer value in two's complement format, with its highest bit being the sign bit and the remaining bits being the numeric bits. This does not change the mathematical meaning of the original FP8 or BP16 exponent, but only its hardware encoding form. At the same time, it can also be synchronously converted to two's complement with the exponent of the fourth floating-point element in the circuit and sent to the exponent comparison module through the same bus structure. This design eliminates the need for additional sign judgment logic in the exponent comparison module, allowing the reuse of standard signed comparator units and reducing the complexity of control logic. In addition, this two's complement representation supports the direct calculation of negative exponent differences, providing a deterministic basis for the subsequent right shift alignment amount (i.e., exponent difference).
[0090] The exponent of the fourth floating-point element is the exponent part of the intermediate result generated by the matrix multiplication and addition module 120 after performing FP8×FP8 or BP16×BP16 multiplication on the first operand and the second operand. The exponent is encoded in two's complement by the exponent adjustment module in the preprocessing module 150 before output; its bit width can be dynamically adapted according to the FP8 or BP16 format type.
[0091] When the input is in E4M3 format, the two's complement exponent has a bit width of 10 bits; when the input is in E5M2 format, the two's complement exponent also has a bit width of 10 bits. Format compatibility can be achieved through high-bit sign extension. The exponent of the fourth floating-point element and the exponent of the floating-point element in the third operand after preprocessing jointly participate in the maximum exponent selection. Their numerical relationship directly determines the number of bits that the mantissa in each path needs to be shifted to the right. At the same time, since both use signed two's complement representation, the exponent difference can be directly output by the subtractor without the need for conditional branches or absolute value correction, which significantly improves the timing robustness of the alignment path.
[0092] Specifically, before the preprocessing module 150 sends the third operand to the accumulation module 130, and before the matrix multiplication and addition module 120 outputs the fourth operand, the preprocessing module 150 performs two's complement encoding on the exponent field: for the original exponent e4∈[0, 15] (with bias) in E4M3 format, it can first be restored to the unbiased exponent e4′=e4-7∈[-6, 8], and then encoded with 10-bit signed two's complement; for the original exponent e5∈[0, 31] (with bias) in E5M2 format, it can first be restored to the unbiased exponent e5′=e5-15∈[-14, 16], and then encoded with 10-bit signed two's complement; the encoding process can be completed collaboratively by the exponent adjustment module in the preprocessing module 150 and the exponent encoding module inside the matrix multiplication and addition module 120. The two modules are symmetrical in structure and aligned in timing, which can ensure that all exponent data entering the accumulation path has a unified two's complement semantics.
[0093] As an example, in an (8×8)×(8×8) matrix multiplication and addition operation, assuming that the first and second operands both use E5M2 format, the exponent of the fourth floating-point element output from a pair of FP8 element multiplications is -17 (corresponding to two's complement 10'b1110110111), while the exponent of the corresponding floating-point element in the preprocessed third operand is -12 (corresponding to two's complement 10'b1111010100). The exponent comparison module performs a signed comparison on these two sets of 10-bit two's complement data, identifies that -12>-17, and therefore selects -12 as the maximum exponent at that position. Subsequently, the alignment and shift module calculates the exponent difference Δ=5 based on this and performs a 5-bit right shift on the mantissa of the fourth floating-point element to complete the alignment. The entire process does not introduce any sign bit discrimination logic or conditional jumps, and is entirely implemented based on the standard two's complement arithmetic unit, meeting the requirements of RISC-V hardware pipelining for low latency and high determinism.
[0094] In this application, the exponents of the floating-point elements in the third and fourth floating-point elements after preprocessing are both represented using signed two's complement. The exponent comparison module can directly perform signed numerical comparisons, avoiding the problem of negative exponents being misjudged as extremely positive numbers under unsigned encoding. Two's complement representation naturally supports negative difference operations, and the alignment shift module can directly output the right shift number using a subtractor, eliminating the need for additional absolute value circuits and multiplexers. The exponents of the two FP8 formats, E4M3 and E5M2, can be mapped to a unified numerical range in the two's complement domain. The same set of exponent alignment hardware can be compatible with dual-format input, improving circuit reuse and architectural simplicity.
[0095] In some embodiments, when the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element reach the minimum negative value represented by the sign complement, independent protection processing is performed on the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element.
[0096] In this embodiment, the exponents of the floating-point elements in the third and fourth floating-point elements after preprocessing are represented using signed two's complement, with a bit width of not less than 10 bits. For example, it can be 10, 11, or 12 bits. The specific bit width can be set by balancing the precision requirements of low-precision operations in RISC-V vector extension with hardware resource constraints.
[0097] When the exponent reaches the smallest negative value represented by the sign complement, it corresponds to the extreme form of all 1s followed by 0s in the complement encoding, such as -512 in 10-bit complement (i.e., binary 1000000000). Its value is usually subjected to conventional arithmetic operations such as negation, left shift (for normalization), addition (for exponent alignment offset), or maximum value filtering by the comparator. It is very easy to cause logical errors or undefined behavior due to overflow.
[0098] The minimum negative value is identified as a boundary exponential state that requires special handling. Its functional meaning is to characterize the deepest underflow exponential level that the FP8 or BP16 denormalized number can reach after preprocessing. It is named based on the fact that it constitutes a systematic risk source in the exponential alignment and normalization link.
[0099] Independent protection processing can be understood as setting up an extreme value detection and response module in the exponential data path that is independent of the main operation path. This module does not participate in regular arithmetic operations, but only performs pattern matching and condition replacement on the input exponential value.
[0100] When the exponents of the third and fourth floating-point elements in the currently processed operand reach the minimum negative value represented by the sign complement, an overflow occurs. This forms an action path with the exponent comparison module and the normalization module. That is, if the input contains the minimum negative value during the process of selecting the maximum exponent, the exponent comparison module must avoid misjudging it as the valid maximum value. If the input contains the minimum negative value during the left / right normalization adjustment, the normalization module must prevent the output of an illegal FP32 exponent due to exponent subtraction overflow. Through the above cooperation, even under the most unfavorable input combination (such as full Denorm input with multiple levels of accumulation), the controllability of the exponent field and the physical consistency of mantissa alignment can be maintained, thereby forming a stable and predictable floating-point operation state flow.
[0101] During the independent protection process of the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element, when it is detected that the exponent of any floating-point element in the preprocessed third operand or the exponent of any floating-point element in the fourth floating-point element is equal to the smallest negative value in the signed two's complement representation, a dedicated protection path is activated. This protection path does not rely on general arithmetic logic units, but is implemented through hard-wired logic.
[0102] Specifically, in the exponent comparison module, the smallest negative value is set as an invalid candidate, forcibly excluding it from participating in the maximum value competition; in the alignment shift module, when the smallest negative value is used as the minuend in the exponent difference calculation, the output is clamped with a fixed safety offset (e.g., 0 or 1), instead of performing actual subtraction; in the normalization module, when the smallest negative value is used as the input exponent in the left normal exponent correction, the exponent update is frozen and a preset compensation value is enabled (e.g., the exponent is forcibly set to -126 to match the lower limit of the FP32 minimum positive normalized number exponent).
[0103] In addition, when the exponent comparison module selects the largest exponent from the exponents of the floating-point elements in the preprocessed third operand and the exponents of multiple fourth floating-point elements, if any input is the smallest negative value, the input is automatically masked by the logic gate circuit, and only the valid exponent participates in the comparison. Subsequently, the alignment and shift module calculates the right shift number of the mantissa of each fourth floating-point element based on the selected largest exponent. At this time, if the exponent of a certain fourth floating-point element was originally the smallest negative value, its right shift number is forced to be set to the maximum allowed right shift amount (e.g., 23 bits), and its mantissa high bits are padded with the sign bit to maintain the numerical meaning. This ensures that even under extreme underflow input conditions, the accumulation operation can still generate an intermediate result with physical interpretability, rather than triggering an abnormal interrupt or uncontrollable output.
[0104] In this application, a protection logic independent of the general arithmetic path is enabled when the minimum negative value occurs, avoiding alignment failure caused by exponent inversion or subtraction overflow. At the same time, the protection logic can form a closed loop with the exponent comparison, alignment shift, and normalization modules, thereby ensuring the numerical robustness of low-precision matrix multiplication and addition operations in the entire input domain. In addition, the protection action can be implemented based on hardware hardwired without software intervention, without introducing additional timing overhead, and maintaining the single-cycle issue characteristic of RISC-V vector instructions.
[0105] In some embodiments, the preprocessing module 150 includes a splitting module, a special data detection module, and a hidden bit recovery module; wherein, the splitting module is connected to the input module 110, the hidden bit recovery module, the matrix multiplication and addition module 120, and the accumulation module 130, respectively; the special data detection module is connected to the input module 110, the hidden bit recovery module, and the accumulation module 130, respectively; and the hidden bit recovery module is connected to the matrix multiplication and addition module 120 and the accumulation module 130, respectively; the splitting module is configured to split the first floating-point element, the second floating-point element, and the third floating-point element into bits respectively to obtain the sign bit, exponent, and mantissa in the first floating-point element, the second floating-point element, and the third floating-point element; the special data detection module is configured to perform normalization detection on the mantissa in the first floating-point element, the second floating-point element, and the third floating-point element respectively to obtain the normalization detection result; and the hidden bit recovery module is configured to recover the hidden bit of the mantissa in the first floating-point element, the second floating-point element, and the third floating-point element based on the normalization detection result.
[0106] In this embodiment, the splitting module can be understood as a logic circuit unit used to perform physical bit field separation on the input FP8 or BP16 format floating-point elements, and is configured to parse each FP8 or BP16 floating-point element into independent sign bit (S), exponent bit (E) and mantissa bit (M) according to a preset format.
[0107] Among them, the FP8 format includes the E4M3 format (1bitS+4bitE+3bitM) and the E5M2 format (1bitS+5bitE+2bitM), and the BP16 format includes 1bitS+8bitE+7bitM.
[0108] The split module can be a combinational logic circuit, without introducing timing delays, and its output signal lines can correspond to three sets of parallel buses: S, E, and M.
[0109] Specifically, the input port of the splitting module interfaces with the 512-bit FP8 / BP16 vector interface of the input module 110, supporting parallel splitting of 64 FP8 elements or 32 BP16 elements simultaneously. The sign bit, exponent bit, and mantissa bit output by the splitting module are respectively transmitted to the special data detection module and the hidden bit recovery module. The function of the splitting module is to provide a structured input foundation for subsequent semantic recognition and numerical reconstruction. It forms the first stage of the preprocessing link with the input module 110 and a data-driven relationship with the special data detection module. That is, only after the bit field splitting is completed can the subsequent modules accurately identify the meaning of each field, so that the FP8 and BP16 dual-format inputs can be uniformly parsed in the same hardware path, avoiding the increase in control complexity caused by the pre-format discrimination.
[0110] The special data detection module can be understood as a dedicated detection circuit used to perform semantic classification on the mantissa obtained by splitting. It is configured to determine the type of the current floating-point element based on the characteristics of the mantissa value and the corresponding exponent state, including five categories: normalized number (Norm), denormalized number (Denorm), zero value (Zero), infinity (Inf), and not-a-number (NaN).
[0111] The detection logic of the special data detection module is based on the IEEE 754 compatibility rules: For the E4M3 format, when the exponent E=0 and the mantissa M=0, it is determined as Zero; when E=0 and M≠0, it is determined as Denorm; when E∈[1, 14], it is determined as Norm; when E=15, it is further distinguished as Inf (M=0) and NaN (M≠0) based on the value of M; for the E5M2 format, when E=0 and M=0, it is determined as Zero; when E=0 and M≠0, it is determined as Denorm. norm; when E∈[1,30], it is determined as Norm; when E=31, it is determined as Inf (M=0) and NaN (M≠0) based on the value of M; for BP16 format, when E=0 and M=0, it is determined as Zero; when E=0 and M≠0, it is determined as Denorm; when E∈[1,254], it is determined as Norm; when E=255, it is determined as Inf (M=0) and NaN (M≠0) based on the value of M.
[0112] The hidden bit recovery module can be understood as a numerical reconstruction unit that pads the hidden bits in the lower bits of the mantissa based on the normalization detection result. It is configured as follows: when the detection result is a normalized number (Norm), a hidden bit of 1 is added before the most significant bit of the mantissa to form a complete normalized mantissa; when the detection result is a denormalized number (Denorm), a hidden bit of 0 is added before the most significant bit of the mantissa to form a complete denormalized mantissa; when the detection result is Zero, Inf, or NaN, the original value of the mantissa is kept unchanged or set to zero.
[0113] The hidden bit recovery module can be composed of a multiplexer and splicing logic. Its inputs include the original mantissa output by the splitting module, the type code output by the special data detection module, and control signals. Its output is the extended mantissa (e.g., E4M3 format extended to 4-bit mantissa, E5M2 format extended to 3-bit mantissa, BP16 format extended to 8-bit mantissa), and the corresponding exponent value is updated synchronously (Denorm needs to correct the exponent according to the number of leading zeros).
[0114] The hidden bit recovery module can provide mantissa input with full precision expression capabilities for subsequent multiplication operations, exponent alignment and normalization processing; together with the special data detection module, it forms the third level of the preprocessing link, realizing the conversion from semantic class to computable value, ensuring that the FP8 or BP16 mantissa has a unified bit width and clear hidden bit semantics before entering the main operation path, avoiding multiplication precision loss or reduction logic error caused by the absence of hidden bits.
[0115] In this application, the splitting module, special data detection module, and hidden bit recovery module constitute a three-stage pipelined preprocessing link. The splitting module first completes the physical decoupling of S / E / M of FP8 or BP16 elements. The special data detection module determines the semantic type based on the joint E and M. The hidden bit recovery module dynamically injects hidden bits according to the type code and outputs the normalized mantissa, so that the same hardware circuit can unambiguously process three low-precision formats: E4M3, E5M2, and BP16. It also provides structurally consistent, semantically clear, and numerically complete input data for the subsequent matrix multiplication and addition module 120 and accumulation module 130. This link does not change the meaning of the original data values, but only completes the format parsing and numerical explicitation. All operations are completed within a single cycle and do not affect the overall pipeline depth.
[0116] In some embodiments, the preprocessing module 150 further includes a leading zero detection module, a mantissa normalization module, and an exponent adjustment module; wherein, the leading zero detection module is connected to the matrix multiplication and addition module 120, the accumulation module 130, the hidden bit recovery module, and the mantissa normalization module, respectively; the mantissa normalization module is connected to the matrix multiplication and addition module 120 and the accumulation module 130, respectively; and the exponent adjustment module is connected to the splitting module, the leading zero detection module, and the matrix multiplication and addition module 120, respectively; the leading zero detection module is configured to perform leading zero detection on the unnormalized mantissas in the first floating-point element, the second floating-point element, and the third floating-point element, respectively, to obtain a first leading zero detection result; the mantissa normalization module is configured to normalize the mantissas in the first floating-point element and the second floating-point element based on the first leading zero detection result; and both the exponent adjustment module and the mantissa normalization module are configured to process the exponent and the unnormalized mantissa in the third floating-point element based on the first leading zero detection result, respectively, to obtain the preprocessed third operand.
[0117] In this embodiment, the leading zero detection module can be understood as a logic unit used to perform a leading zero count (LZC) operation on the denormalized mantissa of the input floating-point element. Its output is a first leading zero detection result representing the number of consecutive zeros in the high bits of the denormalized mantissa. The first leading zero detection result is an unsigned integer, and the bit width is set according to the minimum supported denormalized mantissa length.
[0118] The leading zero detection module is responsible for the identification and quantization of denormalized numbers, and its detection results serve as the control basis for subsequent normalization shift and exponential compensation.
[0119] Meanwhile, the leading zero detection module is connected to the splitting module, hidden bit recovery module, mantissa normalization module, and matrix multiplication and addition module 120 via a multi-channel data bus and control signal line. The first leading zero detection result output by the module is synchronously distributed to the mantissa normalization module and the exponent adjustment module to achieve coordinated response of the three floating-point element processing paths. This allows the denormalized mantissa to complete the valid bit positioning before entering the multiplication operation, avoiding invalid low bits from participating in the operation, thereby reducing power consumption and improving computational efficiency.
[0120] The mantissa normalization module can be understood as performing a left shift alignment operation on the denormalized mantissa based on the first leading zero detection result, and simultaneously updating the combinational logic circuit of the implicit bits of the corresponding floating-point element and the mantissa precision representation.
[0121] The mantissa normalization module can be a pure combinational logic structure or it can include a first-level register for timing constraint optimization. Its inputs include the original mantissa from the splitting module, the normalization status flag from the special data detection module, and the first leading zero detection result from the leading zero detection module. Its output is the normalized mantissa and the corresponding normalization offset.
[0122] The mantissa normalization module can be connected to the leading zero detection module, the matrix multiplication and addition module 120, and the accumulation module 130 via a width-matched data path.
[0123] In this application, the mantissa normalization module is used to normalize the unnormalized mantissas in the first and second floating-point elements, converting them into a normalized representation in the form of 1.xxxx, and explicitly restoring the hidden bits. It works in conjunction with the leading zero detection module to form the core execution unit for the Denorm to Norm conversion, enabling FP8 or BP16 multiplication operations to be performed under the premise of unified normalization, ensuring the consistency of product mantissa precision and the predictability of exponent calculation.
[0124] The exponent adjustment module can be understood as an arithmetic logic unit that performs compensation correction on the exponent field of the third floating-point element based on the first leading zero detection result. The exponent adjustment module can be an adder or a lookup table structure, and its inputs include: the original exponent from the splitting module and the first leading zero detection result from the leading zero detection module; its output is the compensated exponent value.
[0125] The exponent adjustment module can be connected to the splitting module, the leading zero detection module, and the matrix multiplication and addition module 120 via a bus with matching exponent bit width.
[0126] In this application, the exponent adjustment module and the mantissa normalization module work together. For the unnormalized number in the third floating-point element (i.e., the accumulated input C), while normalizing its mantissa, the exponent is adjusted simultaneously to ensure that the exponent reference of the third floating-point element is consistent with the exponent reference of the first and second floating-point elements after normalization. The exponent compensation amount performed by the exponent adjustment module is equal to the number of left shifts indicated by the first leading zero detection result. This ensures that the subsequent exponent alignment module uses the exponent value under the same normalization level when comparing the exponents of multiple floating-point elements, eliminating the exponent alignment deviation caused by the third floating-point element not being synchronously normalized, and improving the accuracy and robustness of numerical alignment in the accumulation stage.
[0127] Specifically, when the input first, second, or third floating-point element is determined to be a denormalized number by the special data detection module, the leading zero detection module immediately performs a leading zero count on its mantissa and outputs the first leading zero detection result. The first leading zero detection result can drive the mantissa normalization module to perform a left shift LZC bit operation on the mantissa of the first and second floating-point elements, and fill the low bits with 0, while restoring the implicit bits to form a complete normalized mantissa. On the other hand, the first leading zero detection result can be synchronously sent to the exponent adjustment module to subtract LZC from the original exponent of the third floating-point element to offset the exponent reduction effect caused by the left shift of the mantissa, thereby maintaining numerical identity. This ensures that all input floating-point elements are in a uniform normalized state and their exponents are comparable before entering the matrix multiplication and addition operation module 120, thus laying the foundation for subsequent high-precision accumulation alignment.
[0128] As an example, before a low-precision matrix multiplication-addition operation is initiated, the input module 110 receives a set of operands in FP8 format. The third operand contains a denormalized number in E4M3 format with a mantissa of 0001000 and an exponent of 0000 (bias 7, actual exponent -7). The leading zero detection module detects that the mantissa has 3 leading zeros and outputs the first leading zero detection result as 3. Based on this, the mantissa normalization module shifts the mantissas of the first and second floating-point elements 3 bits to the left, pads them with zeros to form 10000000, and restores the implicit bit to 1, thus obtaining the normalized number. The mantissa is reduced to 1.0000000; the exponent adjustment module then subtracts 3 from the original exponent 0000 of the third floating-point element, resulting in 1101 (two's complement representation of -3), corresponding to an actual exponent of -10, thus maintaining numerical equivalence with the normalized mantissa 0.0010000×2-7=1.0000000×2-10; after this processing, the third floating-point element is directly comparable to the other two data in the exponent field, and the subsequent exponent alignment module can accurately select the largest exponent for right shift alignment, avoiding alignment errors or underflow misjudgments caused by uncorrected denorm.
[0129] In this application, by setting a leading zero detection module, the valid starting position of the denormalized mantissa can be accurately identified; the mantissa normalization module performs left shift normalization on the mantissa of the first floating-point element and the second floating-point element based on the detection result, which can eliminate redundant low-bit operations when the denormalized number participates in multiplication and improve energy efficiency; the exponent adjustment module simultaneously performs compensation correction on the exponent of the third floating-point element, which can ensure the consistency of the three input data on the exponent reference, ensure the atomicity and timing determinism of the Denorm processing, avoid cross-module asynchronous errors, and thus support the high reliability and low power consumption implementation of FP8 or BP16 to FP32 matrix multiplication and addition operations under the RISC-V architecture.
[0130] In some embodiments, the accumulation module 130 includes an exponent comparison module, an alignment shift module, and an addition module; wherein, the exponent comparison module is connected to the preprocessing module 150, the alignment shift module, and the output module 140, and the addition module is connected to the preprocessing module 150, the alignment shift module, and the output module 140, respectively; the exponent comparison module is configured to select the largest exponent from the exponents of the floating-point elements in the preprocessed third operand and the exponents of multiple fourth floating-point elements to obtain the exponent of the fifth floating-point element; the alignment shift module is configured to right-shift and align the mantissas of multiple fourth floating-point elements based on the largest exponent to obtain multiple right-shifted and aligned mantissas; the addition module is configured to perform an addition operation on the mantissas of the floating-point elements in the preprocessed third operand and the multiple right-shifted and aligned mantissas to obtain the mantissa of the fifth floating-point element.
[0131] In this embodiment, the exponent comparison module can be understood as a dedicated logic unit for performing multi-way floating-point exponent comparison. Its input receives the extended exponents (represented in signed two's complement form) of each floating-point element in the third operand output by the preprocessing module 150, and the extended exponents of multiple fourth floating-point elements output by the matrix multiplication and addition module 120. The module selects the largest exponent from all the aforementioned input exponents through a parallel comparator array or a hierarchical tree comparison structure, and outputs it as the exponent of the corresponding floating-point element in the fifth operand to the alignment shift module and the output module 140. The largest exponent serves as the alignment reference, and its value determines the difference in the number of bits that need to be shifted to the right for all the mantissas involved in the accumulation, thereby ensuring the uniformity of the decimal point position.
[0132] The exponent comparison module supports unified comparison of exponents with different bit widths under both FP8 and BP16 formats. For example, when using FP8 (E4M3), the input exponent is a 4-bit signed two's complement, and when using BP16, the input exponent is an 8-bit signed two's complement. The module achieves compatibility internally through sign bit extension and bit width alignment mechanisms. The output of the exponent comparison module does not introduce new precision loss, but only serves as a control signal to drive subsequent alignment operations. Its delay characteristics are adapted to the RISC-V Vector Extension (RVV) instruction pipeline.
[0133] The alignment shift module can be understood as a configurable shift unit used to perform mantissa right shift alignment. Its input receives the maximum exponent from the exponent comparison module, the third operand mantissa from the preprocessing module 150, and the mantissas of multiple fourth floating-point elements from the matrix multiplication and addition module 120. This module first calculates the exponent difference (i.e., the maximum exponent minus the exponent of the fourth floating-point element itself) corresponding to each fourth floating-point element mantissa, and then performs a logical right shift on the corresponding mantissa based on the difference. During the right shift, a Guard bit, a Round bit, and a Sticky bit (GRS) are generated simultaneously, where the Guard bit is retained after the right shift. The module has the most significant decimal place, the Round bit is the second most significant bit, and the Sticky bit is the logical OR result of the remaining lower bits. This module supports dynamic bit width input, for example, the mantissa is extended to 4 bits (including implicit bits) in FP8 format and 8 bits (including implicit bits) in BP16 format, both of which are uniformly extended to 24 bits to participate in right shift and GRS generation. This module, together with the exponent comparison module, constitutes the alignment stage of floating-point accumulation, ensuring that all mantissas to be accumulated are numerically aligned under the same exponent base, which is a key link to ensure accumulation accuracy. The right-shifted and aligned mantissas output by this module have a uniform decimal point position, providing standardized input for subsequent parallel addition.
[0134] The addition module can be understood as a fixed-point addition unit used to perform parallel reduction addition of multiple operands. Its input receives the mantissa of the third operand after preprocessing, the mantissa of the fourth floating-point element after multiple right shift alignment, and optional guard / round / sticky auxiliary bits.
[0135] The addition module can use a Wallace tree or Han-Carlson addition tree structure to perform parallel compression and final addition on at least 5 input mantissas (8 product mantissas + 1 third operand mantissa in FP8 scenario, a total of 9 paths; 4 product mantissas + 1 third operand mantissa in BP16 scenario, a total of 5 paths) and output a single 24-bit mantissa result.
[0136] Meanwhile, the addition module can also support multi-cycle accumulation mode with carry chain, and can also reduce all input mantissas in a single cycle; its output mantissa is the original mantissa of the fifth operand, which has not yet been normalized, but already has complete precision information.
[0137] In addition, the addition module can form a three-stage pipeline structure of exponent-alignment-addition together with the preprocessing module 150 and the alignment and shifting module, so that the accumulation process is clearly decoupled, the timing is controllable, and it is easy to verify. The mantissa output by the addition module and the maximum exponent output by the exponent comparison module can jointly form the basic elements of the fifth operand, thereby providing a consistent and uniform input basis for the output module 140 to perform leading zero detection and normalization processing.
[0138] The exponent comparison module, alignment shift module, and addition module work collaboratively in sequence according to the data flow: The exponent comparison module first extracts all extended exponents from the preprocessed third operand (i.e., the initial accumulation term) and multiple fourth floating-point elements (i.e., the intermediate product terms generated by matrix multiplication), and quickly determines the maximum exponent value; this maximum exponent is broadcast to the alignment shift module as a unified reference for right-shifting and aligning all mantissas; the alignment shift module calculates the exponent difference of each input mantissa based on this, performs the corresponding right shift operation, and captures the GRS bit for subsequent rounding; the addition module receives all aligned mantissas, performs high-precision fixed-point addition with a unified decimal point position, and outputs the unnormalized fifth floating-point element mantissa; the entire process strictly follows the IEEE 754 style floating-point accumulation principle, but structural optimization and path reuse have been performed for the low-precision scenario of RISC-V to avoid setting up independent accumulation paths for FP8 and BP16 respectively.
[0139] As an example, in an FP8 format operation scenario, the third operand output by the preprocessing module 150 contains one FP8 floating-point element. Each element, after splitting and expanding, has a 4-bit significant mantissa and a 4-bit signed two's complement exponent. The matrix multiplication and addition module 120 outputs eight fourth floating-point elements, each also having a 4-bit significant mantissa and a 4-bit expanded exponent. The exponent comparison module selects the maximum value (e.g., +3) from all 16 exponents and sends this value to the alignment and shift module. The alignment and shift module calculates the difference between the exponent of each fourth floating-point element and +3 (e.g., if the exponent of an element is +1, the difference is 2), performs a 2-bit logical right shift on the mantissa of that element, and generates the corresponding GRS bit. The addition module receives the mantissa of the third operand and the eight right-shifted and aligned mantissas (a total of nine inputs), compresses them using a Wallace tree, and outputs a single 24-bit mantissa result, which is the mantissa at the corresponding position in the fifth operand. This process is completed in a single cycle, meeting the throughput requirements of RISC-V vector instructions.
[0140] In this application, the exponent comparison module uniformly selects the largest exponent from the preprocessed third operand and multiple fourth floating-point elements, thus solving the problem of accumulation precision loss caused by inconsistent exponents of multi-source floating-point data; the alignment and shift module performs right shift alignment on the mantissa of all fourth floating-point elements based on the largest exponent, ensuring that the decimal point position of different exponent data is unified before addition, improving the consistency of accumulated values; the addition module participates in parallel addition with the mantissa of the preprocessed third operand and the mantissa of the right-shifted and aligned fourth floating-point elements, supporting flexible adaptation to different scale inputs such as FP8 (8-way product + 1-way accumulation) and BP16 (4-way product + 1-way accumulation), enhancing the circuit's compatibility and configurability with the RISC-V vector extension instruction set.
[0141] In some embodiments, the output module 140 is configured to perform leading zero detection on the mantissas of a plurality of fifth floating-point elements to obtain a second leading zero detection result, and to normalize the mantissas of the plurality of fifth floating-point elements based on the second leading zero detection result to generate the mantissa in the target operand; to adjust the exponents of the plurality of fifth floating-point elements based on the second leading zero detection result to generate the exponent in the target operand; and to output the target operand.
[0142] The output module 140 can be used to perform final formatting on the fifth floating-point element formed after the accumulation operation. The output module 140 can convert the denormalized mantissa into the 1.xxxx normalized form required by the IEEE 754 FP32 standard and simultaneously correct the corresponding exponent, thereby ensuring that the output result meets the RISC-V floating-point specification requirements in terms of numerical representation, rounding behavior and exception flag generation.
[0143] The output module 140 works in conjunction with the exponent comparison module to receive the maximum exponent value of its output as a normalization reference. At the same time, the output module 140 can also work in conjunction with the addition module to receive the unnormalized mantissa and sign bit of its output, forming a complete input data stream. This enables integrated processing of mantissa left / right shift, exponent increment / decrement, GRS bit extraction and rounding decision, forming a closed-loop path from the accumulation result to the standard FP32 output.
[0144] Specifically, in the process of performing leading zero detection on the mantissa of multiple fifth floating-point elements to obtain the second leading zero detection result, this application can use a leading zero counter (LZC) circuit to perform parallel leading zero statistics on the mantissa of each fifth floating-point element output by the addition module to output the second leading zero detection result, which represents the number of consecutive zeros starting from the most significant bit.
[0145] The second leading zero detection result can be an integer between 0 and 23, where 0 indicates that the highest bit of the mantissa is 1 (normalized), and 23 indicates that the mantissa is all 0 (i.e., the result is zero).
[0146] Meanwhile, in the process of performing leading zero detection on the mantissas of multiple fifth floating-point elements to obtain the second leading zero detection result, this application can be executed based on the original code mantissa, or the mantissa in two's complement can be uniformly converted to the original code form after sign extension before execution; the second leading zero detection result can be directly used for subsequent normalized shift amount determination and GRS bit generation.
[0147] For example: when the second leading zero detection result is n, shifting left by n bits can align the most significant bit to the 23rd bit (the most implicit bit of the FP32 mantissa); when the second leading zero detection result is 0 and the two most significant bits of the mantissa are 10, it indicates that there is a carry overflow, and it is necessary to right-align by 1 bit and increment the exponent by 1.
[0148] In the process of normalizing the mantissas of multiple fifth floating-point elements based on the second leading zero detection result to generate the mantissas in the target operand, the programmable shifter can be controlled to perform left or right shift operations on the mantissas according to the second leading zero detection result.
[0149] When the second leading zero detection result n>0, a left shift of n bits is performed, with low-order bits padded with 0; when the second leading zero detection result n=0 and the high-order bit of the mantissa is 10, a right shift of 1 bit is performed, with high-order bits discarded and low-order bits padded with rounding bits; the shifted mantissa retains 23 significant decimal places and reserves 1 implicit bit space; this normalization process can be completed in conjunction with rounding operations, for example, during the left shift, low-order redundant bits are truncated simultaneously and three rounding auxiliary bits G / R / S are generated; or after the shift is completed, an independent rounding module generates the final Rnd bit based on the GRS bit and RISC-V rounding mode (RNE / RTZ / RDN / RUP / RMM) and performs addition; the normalized mantissa is a standard FP32 format 23-bit explicit mantissa, with its most significant bit always 1 (implicit) or all zeros (zero value output).
[0150] In the process of adjusting the exponents of multiple fifth floating-point elements based on the second leading zero detection result to generate the exponent in the target operand, the exponents of multiple fifth floating-point elements (denoted as Exp_max) output by the exponent comparison module can be arithmetically operated with the second leading zero detection result n.
[0151] When left-shifting by n bits, the exponent is updated to Exp_out = Exp_max - n; when right-shifting by 1 bit, the exponent is updated to Exp_out = Exp_max + 1. This exponent adjustment process is implemented using a signed adder, supporting a 10-bit signed exponent operation range [-162, 386]. The adjusted exponent is mapped to the FP32 exponent field [1, 254] after being limited: if Exp_out > 254, it is determined to be an overflow (OF), and Inf or a saturation value is output; if Exp_out < 1 and the mantissa is not zero, it is determined to be an underflow (UF), and a denormalized number or zero is output. This mapping process is completed synchronously with the exception flag generation logic.
[0152] Finally, during the output of the target operand, the normalized mantissa, adjusted exponent, and original sign bit can be packaged and output in FP32 format. The sign bit comes directly from the sign bit output by the addition module, the exponent is an 8-bit unsigned integer (bias 127), and the mantissa is a 23-bit explicit decimal. The output data is written to the vector register file in cycles via a 512-bit bus, with 8 FP32 elements output per cycle, for a total of 4 cycles to complete the output of 64 FP32 results. At the same time, the output process is controlled by the RISC-V vector instruction vfmopma.vv, supporting mask writing and broadcast modes.
[0153] The output module 140 operates as follows: it receives the unnormalized mantissa and sign bit from the addition module and the maximum exponent value from the exponent comparison module; it calls the LZC circuit to perform leading zero detection on each mantissa to obtain the second leading zero detection result; based on the result, it determines whether to shift left, shift right, or keep it unchanged, and drives the shifter to perform the corresponding operation; synchronously, it sends the second leading zero detection result to the exponent adjustment unit to calculate the output exponent; at the same time, during or after the shift, it generates three bits (G / R / S) based on the low-order truncation information of the mantissa and sends them to the rounding module to participate in the Rnd bit calculation; finally, it combines the sign bit, the adjusted exponent, the normalized mantissa, and the exception flag into a standard FP32 word, which is then driven to the external interface via the output buffer.
[0154] As an example, after performing matrix multiplication and addition from (8×8)×(8×8) to (8×8), the accumulation module 130 outputs 64 fifth floating-point elements, each containing a 32-bit mantissa field (including implicit bits), a 10-bit exponent field, and a 1-bit sign bit. The output module 140 performs LZC detection on all 64 mantissas in parallel, obtaining 64 sets of second leading zero detection results. For one set of mantissas with a detection result of 3, a left shift operation of 3 bits is performed, changing the original mantissa 0001_0101_… to 1010_1000_…, while decrementing the exponent by 3. For another set of mantissas with a detection result of 0 and a high-order bit of 10, a right shift operation of 1 bit is performed, changing the original mantissa 1011_… to 0101_…, while incrementing the exponent by 1. All shifted mantissas are processed by the rounding module, adding Rnd bits according to the RNE mode and truncating to 23 bits. Finally, 64 elements conforming to IEEE 754 are output. The target operand in FP32 format is written to four vector registers dst0–dst3, each register holding 16 FP32 elements.
[0155] In this application, the output module 140 performs normalization processing on the mantissa of the fifth floating-point element based on the second leading zero detection result, thereby restoring the denormalized mantissa to the standard form of 1.xxxx; the corresponding exponent is adjusted synchronously to ensure that the numerical precision is not lost due to shift; the normalization process is deeply coupled with GRS bit generation and rounding mode selection, and can support all five rounding behaviors defined by RISC-V; the exponent adjustment range covers [-162, 386] and overflow / underflow criteria are set, which can accurately generate abnormal flags such as OF / UF / NX, so that the final state normalized output of the FP8 matrix multiplication and addition link can be completed without introducing a new structure.
[0156] In some embodiments, the matrix multiplication and addition module 120 includes a multiplexer and a multiplication module; wherein the multiplexer is connected to the input module 110 and the multiplication module, and the multiplication module is connected to the accumulation module 130; the multiplexer is configured to periodically output multiple first floating-point elements of a preset number of rows in the first operand based on identification information, and the multiplication module is configured to perform row and column multiplication operations on the multiple first floating-point elements of the preset number of rows in the first operand and the second floating-point elements in the second operand in the form of a matrix to obtain multiple fourth floating-point elements.
[0157] In this embodiment, the multiplexer can be understood as a hardware selection unit used to periodically schedule the first operand data stream. Its input port is connected to multiple sets of first floating-point elements output by the input module 110, and its output port is connected to the multiplicand input of the multiplication operation module. The multiplexer is configured to periodically select multiple first floating-point elements of a preset number of rows (e.g., 2 rows, 4 rows, or 8 rows) from the first operands cached by the input module 110 according to the row index required for the current matrix block calculation, and output the selected row data to the multiplication operation module in parallel in each clock cycle. The control signal of the multiplexer comes from the matrix size configuration field generated after RISC-V instruction decoding and the current calculation step counter, and its switching behavior matches the pipeline depth of the matrix multiplication and addition operation.
[0158] In this application, the multiplexer can perform lossless routing selection without changing the numerical format and bit width of the first floating-point element. Its selection path can be configured as static routing or dynamic decoding control according to actual deployment requirements.
[0159] In FP8 operation mode, the multiplexer outputs 8 first floating-point elements in FP8 format to form a row each time; in BP16 operation mode, it outputs 4 first floating-point elements in BP16 format to form a row each time; the multiplexer and the input module 110 work together to implement row-major access scheduling of the first operand, providing structured data supply for subsequent row and column multiplications.
[0160] The multiplication module can be understood as a dedicated computing unit that performs low-precision floating-point multiplication. Its multiplicand input is connected to the output of a multiplexer, and its multiplier input is connected to the second operand output by the input module 110. The multiplication module is configured to perform parallel multiplication operations on multiple first floating-point elements in a preset row number output by the multiplexer and multiple second floating-point elements in the corresponding column of the second operand, in a matrix multiplication row-column correspondence relationship.
[0161] For example, in FP8×FP8 operation, when the multiplexer outputs a line of 8 FP8 elements (denoted as Ai, 0~Ai, 7), and the second operand is an 8×8 matrix, the multiplication module simultaneously calculates Ai, j×Bj, k (j=0~7), generating a total of 8 intermediate products, each containing 8 fourth floating-point elements.
[0162] The multiplication module supports both FP8 and BP16 formats. Its internal multiplier bit width is configurable: in FP8 mode, it adopts a 4×4 Booth-4 encoder + Wallace tree compression structure, and outputs a fourth floating-point element with an extended mantissa of 24 bits and an extended exponent of 10 bits.
[0163] In BP16 mode, an 8×8 Booth-4 encoder + Wallace tree compression structure can be used to output the sixth floating-point element in the format of Sgn(1) + Exp(10) + Man(24).
[0164] The exponent calculation logic of the multiplication operation module is E6=EA+EB-bias-lzcA-lzcB, where bias is dynamically selected according to the format (7 or 15 for FP8, 127 for BP16), and lzc is the leading zero detection result.
[0165] The multiplication module, in conjunction with the multiplexer, enables vectorized loading of the first operand and broadcast column matching of the second operand. It is the core execution unit for performing row and column multiplication operations in matrix form.
[0166] The multiplexer and multiplication module constitute a complete matrix multiplication and addition data path: the multiplexer periodically outputs a row of data from the first operand according to the instruction configuration; the multiplication module multiplies this row of data element-wise with each column of data from the second operand, generating a fourth floating-point element array of the corresponding column dimension; this process is repeated in each calculation cycle until all rows of the first operand and all columns of the second operand are covered, thus completing the entire matrix multiplication operation; it can support matrix block calculation of different sizes, for example, when the first operand is 8×4 and the second operand is 4×8, the multiplexer outputs 2 rows of data 4 times, each time driving the multiplication module to complete the multiplication of the 2×4 and 4×8 submatrixes; the entire process does not require software intervention for scheduling and is fully implemented by hardware pipeline control.
[0167] As an example, when performing a BP16×BP16 to FP32 matrix multiplication and addition operation, the input module 110 provides the first operand (8×4 BP16 matrix) to the multiplexer. The multiplexer outputs the 8 BP16 elements of rows 0–1 at a 2-row / cycle rhythm. At the same time, the input module 110 transposes the second operand (4×8 BP16 matrix) to 8×4 and sends it to the multiplication module. The multiplication module maps each group of 2×4 and 8×4 data into 4 groups of parallel 2×8 multiplications. Each group generates 2×8 fourth floating-point elements (Sgn1+Exp10+Man24), which are sent to the accumulation module 130 and the third operand (FP32 format) to complete the subsequent accumulation, so as to complete the BP16×BP16 to FP32 matrix multiplication and addition operation within 4 cycles.
[0168] In this application, the multiplexer is configured to periodically output multiple first floating-point elements of a preset number of rows in the first operand based on identification information, which enables block scheduling and pipelined loading of large-scale matrices and solves the technical problem that large-size matrices cannot be loaded into the computing unit at one time. The multiplication module is configured to perform row and column multiplication operations on multiple first floating-point elements of a preset number of rows in the first operand and second floating-point elements in the second operand in matrix form. The same set of multipliers can be reused to support both FP8 and BP16 formats, avoiding hardware redundancy caused by designing separate multiplication paths for different precisions.
[0169] In some embodiments, the matrix multiplication and addition module 120 further includes a processing module; wherein the processing module is connected to the accumulation module 130; the multiplication module is configured to perform row and column multiplication operations on a plurality of first floating-point elements in a preset number of rows in the first operand and second floating-point elements in the second operand in the form of a matrix to obtain a plurality of sixth floating-point elements; the processing module is configured to extend the mantissas of the plurality of sixth floating-point elements and perform an alignment operation on the exponents of the plurality of sixth floating-point elements to obtain a fourth operand.
[0170] In this embodiment, the sixth floating-point element is an intermediate result output by the multiplication module after performing a matrix multiplication operation on multiple first floating-point elements in a preset number of rows in the first operand and second floating-point elements in the second operand.
[0171] The sixth floating-point element can serve as a transition point from FP8 or BP16 multiplication precision to FP32 accumulation precision. It forms a format adaptation front-end with the processing module and achieves low-latency data flow through a direct signal path. As a result, the numerical integrity of the sixth floating-point element can be calibrated before entering the accumulation stage, thereby avoiding precision loss caused by mantissa truncation or exponent misalignment.
[0172] The processing module is configured to extend the mantissa of multiple sixth-point elements, thereby extending the mantissa of the sixth-point elements to 24 bits. The extension method can be to pad with zeros at the lower bits, specifically by appending multiple zero bits after the least significant bit of the mantissa, to reserve the precision margin required for subsequent rounding and normalization.
[0173] Meanwhile, the extended mantissa can participate in subsequent exponent alignment and right shift alignment, thus ensuring that the mantissa precision of the final fifth floating-point element meets the requirement of 23 valid mantissas in the FP32 format.
[0174] During the process of aligning the exponents in multiple sixth-point elements, the processing module can uniformly convert the extended exponents carried by different sixth-point elements into signed two's complement representations with the same bit width (e.g., 10 bits), and adapt their numerical range to the exponent comparison requirements of the FP32 accumulation stage.
[0175] The exponent alignment operation can be as follows: when the sixth floating-point element comes from E4M3 format input, its original 4-bit exponent is mapped to the 10-bit two's complement field after offset correction (-7) and LZC adjustment; when it comes from E5M2 format input, its original 5-bit exponent is mapped to the same 10-bit two's complement field after offset correction (-15) and LZC adjustment.
[0176] Specifically, the exponent alignment operation can be performed in conjunction with the mantissa expansion operation. Their connection is achieved through timing synchronization via shared control signals, which in turn provides a directly comparable exponent input to the exponent comparison module in the accumulation module 130. At the same time, the aligned exponent can participate in subsequent maximum exponent filtering and mantissa right shift alignment, which can ensure that multiple sixth floating-point elements and third floating-point elements are comparable and operable in the exponent dimension.
[0177] Specifically, during the process of mantissa expansion and exponent alignment of the sixth floating-point element, the processing module can output 8 groups of the sixth floating-point element in each calculation cycle. Each group contains 1 sign bit, 10 two's complement exponent, and 16 mantissa. After receiving the 8 groups of data, the processing module first performs low-bit zero-padding expansion on the mantissa of each group to generate 8 groups of 24-bit mantissa. At the same time, it performs sign bit expansion and field mapping verification on the 10-bit two's complement exponent of each group to ensure that its value is within the range of [-162, 386] and there is no overflow. Then, the expanded 24-bit mantissa and the aligned 10-bit exponent are packaged into a fourth operand of a unified format and sent to the accumulation operation module 130. This process does not introduce new functional units or control logic, but only relies on the existing bit width expander and exponent remapping circuit. Its hardware overhead is controllable and the timing path is stable.
[0178] As an example, taking the (8×8) matrix multiplication and addition in FP8 format as an example, in the first calculation cycle, the processing module immediately performs mantissa expansion to 24 bits on the 8 sixth floating-point elements output by the multiplication module, and performs normalization alignment on their 10-bit two's complement exponents—unifying the exponent values [-11, 23] generated by the E4M3 path and the exponent values [-17, 47] generated by the E5M2 path to the effective subset of the 10-bit two's complement full range [-512, 511]. The 8 sets of data after expansion and alignment constitute the fourth operand, with a mantissa width of 24 bits, an exponent width of 10 bits, and a sign bit of 1 bit, which fully matches the interface requirements of the accumulation module 130 for the input data format, thereby ensuring that the subsequent accumulation calculation with the third floating-point element (FP32 format) is consistent in both numerical precision and dynamic range.
[0179] As an example, taking the (8×4)×(4×8)→(8×8) matrix multiplication and addition scenario in BP16 format as an example, 16 groups of sixth floating-point elements are generated in parallel reduction each cycle (each group corresponds to an FP32 output position). The processing module performs mantissa expansion and exponent alignment on the four sixth floating-point elements in each group. For example, when the exponents of the four seventh floating-point elements in a group are 102, 105, 103, and 105 respectively, the largest exponent of 105 is taken as the reference, and the mantissas corresponding to exponents 102 and 103 are right-shifted by 3 bits and 2 bits respectively, and the corresponding GRS bits are generated. After the mantissa is right-shifted, it is uniformly expanded to 24 bits, which together with the aligned exponents form a well-formatted fourth operand, which is used by the accumulation module to perform high-precision accumulation with the FP32 initial value.
[0180] In this application, a processing module is added to the matrix multiplication and addition module 120 to perform mantissa expansion and exponent alignment operations on the seventh floating-point element output by the multiplication module, thereby eliminating the format gap between the multiplication result and the accumulated input. The mantissa expansion is 24 bits with zero padding at the low bits, which provides sufficient precision margin for subsequent addition operations and avoids the loss of low-bit information due to insufficient bit width.
[0181] In some embodiments, this application also provides a processor, including the matrix multiplication and addition circuit based on the RISC-V architecture provided in this application.
[0182] In this application, the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0183] In some embodiments, this application also provides an electronic device, including a matrix multiplication and addition circuit based on the RISC-V architecture provided in this application or a processor provided in this application.
[0184] In this embodiment, the electronic device may be equipped with a processor, which may contain a matrix multiplication and addition circuit based on the RISC-V architecture and be connected to the system bus to provide computing and control capabilities to support the operation of the electronic device.
[0185] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A matrix multiplication and addition circuit based on RISC-V architecture, characterized in that, include: An input module is configured to input operands and perform precision identification on the floating-point elements in the input operands to obtain identification information; the operands include a first operand, a second operand, and a third operand, and the third operand includes multiple third floating-point elements in a preset format; A matrix multiplication and addition module, connected to the input module, is configured to periodically perform row and column multiplication operations on multiple first floating-point elements in a preset number of rows of the first operand and second floating-point elements in the second operand in matrix form based on the recognition information to obtain a fourth operand; the fourth operand includes multiple fourth floating-point elements in a preset format; The first floating-point element and the second floating-point element are both floating-point numbers in FP8 format, or the first floating-point element and the second floating-point element are both floating-point numbers in BP16 format. An accumulation operation module is connected to the matrix multiplication and addition operation module and the input module, and is configured to accumulate multiple fourth floating-point elements with the third floating-point elements to obtain a fifth operand, wherein the fifth operand includes multiple fifth floating-point elements in a preset format; The output module is connected to the accumulation operation module and is configured to output the target operand based on multiple fifth floating-point elements; A preprocessing module, wherein the input end of the preprocessing module is connected to the input module, and the output end of the preprocessing module is connected to the matrix multiplication and addition operation module; The preprocessing module is configured to preprocess the first operand, the second operand, and the third operand respectively to obtain the preprocessed first operand, the second operand, and the third operand. The exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element are represented using signed two's complement. When the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element reach the minimum negative value represented by signed two's complement, independent protection processing is performed on the exponents of the floating-point elements in the preprocessed third operand and the fourth floating-point element.
2. The matrix multiplication and addition circuit based on RISC-V architecture according to claim 1, characterized in that, The preprocessing module includes a splitting module, a special data detection module, and a hidden bit recovery module; The splitting module is connected to the input module, the hidden bit recovery module, the matrix multiplication and addition module, and the accumulation module, respectively. The special data detection module is connected to the input module, the hidden bit recovery module, and the accumulation module, respectively. The hidden bit recovery module is connected to the matrix multiplication and addition module and the accumulation module, respectively. The splitting module is configured to split the first floating-point element, the second floating-point element, and the third floating-point element into bits respectively, so as to obtain the sign bit, exponent, and mantissa in the first floating-point element, the second floating-point element, and the third floating-point element. The special data detection module is configured to perform normalization detection on the mantissas of the first floating-point element, the second floating-point element, and the third floating-point element respectively, and obtain normalization detection results. The hidden bit recovery module is configured to recover the hidden bit of the mantissa in the first floating-point element, the second floating-point element, and the third floating-point element based on the normalization detection result.
3. The matrix multiplication and addition circuit based on RISC-V architecture according to claim 2, characterized in that, The preprocessing module also includes a leading zero detection module, a mantissa normalization module, and an exponent adjustment module. The leading zero detection module is connected to the matrix multiplication and addition module, the accumulation module, the hidden bit recovery module, and the mantissa normalization module, respectively. The mantissa normalization module is connected to the matrix multiplication and addition module and the accumulation module, respectively. The exponent adjustment module is connected to the splitting module, the leading zero detection module, and the matrix multiplication and addition module, respectively. The leading zero detection module is configured to perform leading zero detection on the denormalized mantissas in the first floating-point element, the second floating-point element, and the third floating-point element respectively, to obtain the first leading zero detection result; The mantissa normalization module is configured to normalize the unnormalized mantissas in the first floating-point element and the second floating-point element based on the first leading zero detection result. Both the exponent adjustment module and the mantissa normalization module are configured to process the exponent and the denormalized mantissa in the third floating-point element based on the first leading zero detection result, so as to obtain the preprocessed third operand.
4. The matrix multiplication and addition circuit based on RISC-V architecture according to claim 1, characterized in that, The accumulation operation module includes an exponent comparison module, an alignment shift module, and an addition operation module; The exponent comparison module is connected to the preprocessing module, the alignment shift module, and the output module, respectively, and the addition module is connected to the preprocessing module, the alignment shift module, and the output module, respectively. The exponent comparison module is configured to filter out the largest exponent from the exponents of the floating-point elements in the preprocessed third operand and the exponents of the multiple fourth floating-point elements to obtain the exponent of the fifth floating-point element. The alignment shift module is configured to right-shift and align the mantissas of multiple fourth floating-point elements based on the maximum exponent, thereby obtaining multiple right-shifted and aligned mantissas. The addition module is configured to perform an addition operation on the mantissa of the floating-point element in the preprocessed third operand and multiple right-shifted mantissas to obtain the mantissa of the fifth floating-point element.
5. The matrix multiplication and addition circuit based on RISC-V architecture according to claim 4, characterized in that, The output module is configured to perform leading zero detection on the mantissas of the plurality of fifth floating-point elements to obtain a second leading zero detection result, and to normalize the mantissas of the plurality of fifth floating-point elements based on the second leading zero detection result to generate the mantissas in the target operand. Based on the second leading zero detection result, the exponents of multiple fifth floating-point elements are adjusted to generate the exponent in the target operand; Output the target operand.
6. The matrix multiplication and addition circuit based on RISC-V architecture according to claim 1, characterized in that, The matrix multiplication and addition module includes a multiplexer and a multiplication module; The multiplexer is connected to the input module and the multiplication module, respectively, and the multiplication module is connected to the accumulation module. The multiplexer is configured to periodically output a plurality of the first floating-point elements in the first operand, based on the identification information; The multiplication module is configured to perform row and column multiplication operations on multiple first floating-point elements in a preset number of rows in the first operand and second floating-point elements in the second operand in the form of a matrix, to obtain multiple fourth floating-point elements.
7. The matrix multiplication and addition circuit based on RISC-V architecture according to claim 6, characterized in that, The matrix multiplication and addition module also includes a processing module; The processing module is connected to the accumulation operation module; The multiplication module is configured to perform row and column multiplication on multiple first floating-point elements in a preset row number in the first operand and second floating-point elements in the second operand in the form of a matrix to obtain multiple sixth floating-point elements. The processing module is configured to expand the mantissas of the plurality of sixth floating-point elements and align the exponents of the plurality of sixth floating-point elements to obtain the fourth operand.
Citation Information
Patent Citations
Apparatus, method and system for 8-bit floating point matrix dot product instructions
CN118605946A
Multi-precision matrix calculation unit and use method thereof
CN121167096A
Floating point multiply-accumulate unit facilitating variable data precision
CN121420281A