FPGA multiplier optimization method and device
By detecting the effective bit width of the input data and selecting the appropriate calculation mode, resource multiplexing and parallel computing of the FPGA multiplier are realized, which solves the problems of resource waste and low efficiency of traditional multipliers in low bit width scenarios, and improves the computing efficiency.
Patent Information
- Application Number
- CN202510855502.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
When traditional FPGA multipliers process input data with actual bit widths less than 18 bits, there are problems of wasted hardware resources and low computing efficiency, especially in audio processing, neural networks and embedded computing tasks.
The leading zero circuit detects the effective bit width of the input data, generates a bit width control signal, selects the small bit width parallel calculation mode or the full bit width calculation mode, and performs sign bit expansion and data splicing in the small bit width mode to realize parallel operation; directly calculates in the full bit width mode.
The space utilization and computing efficiency of the multiplier are improved, especially in low-bit wide and high-frequency computing scenarios, which significantly improve hardware resource utilization and reduce hardware redundancy.
Smart Images

Figure CN120353430A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of digital integrated circuits, and particularly relates to an FPGA multiplier optimization method and device. Background Art
[0002] With the rapid progress of modern electronic technology and digital signal processing (DSP) technology, DSP modules in FPGAs have been widely used in the field of signal operation and processing. These advanced modules can efficiently implement complex algorithms, providing flexible and powerful solutions to meet the growing high-performance computing needs. By utilizing the DSP resources integrated in FPGAs, engineers can design more efficient and accurate digital signal processing systems, which are applicable to multiple industries such as communication, audio processing, and image processing.
[0003] The multiplier in the DSP module is a core component. The traditional 18bit×18bit fixed-bitwidth multiplier has high efficiency in hardware implementation. However, for some specific application scenarios, such as audio processing, neural networks, and embedded computing tasks, the actual bitwidth of the input data is often less than 18bit and has a large randomness. In this case, the unused high bits result in waste of hardware resources. The traditional multiplier can only process one set of multiplication operations at a time and cannot utilize the remaining idle bitwidth resources for parallel computing, reducing the computing efficiency. Summary of the Invention
[0004] The present invention provides an FPGA multiplier optimization method and device, which combines the input data with a relatively small actual bitwidth in the multiplier to improve the utilization efficiency of the multiplier space.
[0005] Other objects and advantages of the present invention can be further understood from the technical features disclosed in the present invention.
[0006] To achieve one or part or all of the above objects or other objects, an FPGA multiplier optimization method provided by a technical solution of the present invention detects the actual effective bitwidth of the data at two input ends, generates a bitwidth control signal, and based on the bitwidth control signal, selects a small-bitwidth parallel computing mode or a full-bitwidth computing mode; in the small-bitwidth parallel computing mode, the input data with a small effective bitwidth in multiple groups is subjected to sign-bit extension to generate parallel operation data including sign bits, and the generated parallel operation data is subjected to a shift operation and splicing and then sent into the multiplier for parallel operation; in the full-bitwidth computing mode, the data directly enters the multiplier for operation.
[0007] Use a leading zero circuit to determine the number of leading zeros in the input data, identify the highest non-zero bit in the input data, and then output a leading zero detection result to represent the effective bit width of the input data; generate a bit width control signal based on the leading zero detection result; select a corresponding calculation mode according to the bit width control signal and the set small bit width data merging threshold.
[0008] The leading zero circuit captures the input data at the active edge of the clock signal, scans bit by bit from the highest bit to the lowest bit of the data, counts the consecutive zeros until the first 1 is counted, and outputs the leading zero detection result.
[0009] The bit width control signal is obtained by subtracting the leading zero detection result from the multiplier bit width of the data input; if the bit width control signals of the two input terminal data are both less than or equal to the small bit width data merging threshold, start the small bit width parallel calculation mode; if any one of the bit width control signals of the two input terminal data is greater than the small bit width data merging threshold, start the full bit width calculation mode.
[0010] The multiplier bit width is 18, and the small bit width data merging threshold is 8.
[0011] When selecting the small bit width parallel calculation mode, composite the small bit width valid bit data, including: perform bit width detection and sign bit extension processing on the input data to obtain the valid bit data including the sign bit, perform a shift operation on the valid bit data after sign bit extension processing, including based on the actual bit width of the detected input data and the preset shift amount of each group of data, perform a shift operation on the valid bit data after sign bit extension processing, after the shift operation, obtain multiple groups of small bit width operation data, and then splice and combine the multiple groups of small bit width operation data together through a logical "OR" to align to a specific segment of the multiplier; the spliced multiple groups of the small bit width operation data are sent to the multiplier, and each group of small bit width operation data is independently operated.
[0012] The bit width control signal determines the shift amount of the input data.
[0013] The output result of the multiplier is divided into multiple segments, and each segment output result corresponds to the operation result of a group of small bit width operation data; the shifter extracts each segment operation result to obtain the operation results of each group of small bit width valid bit data.
[0014] Before splicing the small bit width operation data, use dynamic control logic to monitor the total bit width of the spliced multiple groups of small bit width operation data, including calculating the total bit width of the currently spliced multiple groups of small bit width operation data, and before each splicing, judge the relationship between the remaining bit width of the multiplier input data and the bit width of the small bit width operation data to be spliced. When the remaining bit width of the multiplier input data is greater than or equal to the bit width of the small bit width operation data to be spliced, continue the data splicing, otherwise stop splicing and send the spliced data to the subsequent module for calculation.
[0015] After multiple groups of small-bitwidth operation data are spliced, starting from the low bit, they are arranged bit by bit in sequence to the input end of the multiplier, and the highest empty bit is filled with 0.
[0016] An FPGA multiplier optimization device provided by another technical solution of the present invention is used to implement an FPGA multiplier optimization method described above. The multiplier optimization device includes: an input bitwidth detection module that uses a leading zero detection circuit to detect the bitwidth of the valid bits of the input data and generates a bitwidth control signal. Based on the bitwidth control signal, a small-bitwidth parallel calculation mode or a full-bitwidth calculation mode is selected; a bitwidth detection and sign bit extension module that, in the case of starting the small-bitwidth parallel calculation mode, performs bitwidth detection and sign bit extension processing on the input data. The valid bit data after the sign bit extension processing is subjected to a shift operation, including shifting the valid bit data after the sign bit extension processing based on the actual bitwidth of the detected input data and the preset shift amount for each group of data. After the shift operation, multiple groups of small-bitwidth operation data are obtained; a splicing judgment module that judges the relationship between the remaining bitwidth of the multiplier input data and the bitwidth of the multiple groups of small-bitwidth operation data to be spliced and decides whether to continue data splicing; a bitwidth splicing and multiplexing module that, in the small-bitwidth parallel calculation mode, splices the small-bitwidth operation data to be spliced, and after the multiple groups of small-bitwidth operation data are spliced, starting from the low bit, they are arranged bit by bit in sequence to the input end of the multiplier, and the empty bits at the high bit are filled with 0; a parallel multiplication calculation module that performs parallel calculation on the spliced small-bitwidth operation data input and outputs multiple segments of operation results, and each segment of output result corresponds to the operation result of a group of small-bitwidth operation data; a shifter extracts each segment of operation result to obtain the operation results of the valid bits of each group of small-bitwidth data; when the full-bitwidth calculation mode is enabled, the data is directly input into the multiplier for operation to output the multiplication result.
[0017] Compared with the prior art, the beneficial effects of the present invention mainly include: 1. The method for realizing adaptive data bitwidth detection by using the leading zero circuit logic in the present invention can dynamically identify the valid bitwidth of the input data and adjust the data splicing and calculation strategies according to the detection results. When the bitwidth is from 1 to 8 bits, resource multiplexing is achieved through splicing; when the bitwidth is from 9 to 18 bits, it is directly input into the multiplier for calculation without additional operations. This characteristic enables the present invention to flexibly adapt to various bitwidth scenarios and meet different precision requirements.
[0018] 2. The data bit-width splicing and multiplexing mode adopted by the present invention greatly improves the utilization efficiency of bit-width resources, can fully utilize the input resources of the multiplier. When performing small bit-width multiplication operations (such as 1×1, 2×2, 3×3, …, 8×8), it supports bitwise splicing of multiple small bit-width data to the input end of the multiplier, so as to realize multiple groups of small bit-width calculations in one operation. This feature significantly improves the operation ability of the multiplier, especially in the low bit-width and high-frequency calculation scenarios (such as audio processing and image processing), greatly improving the utilization rate of hardware resources and operation efficiency, and reducing hardware redundancy.
[0019] 3. The FPGA multiplier optimization device of the present invention has the advantages of simplicity and strong scalability. The modules such as adaptive detection, shifting, splicing, and multiplexer adopted by the present invention realize the functions of bit-width adaptive detection and multiplexing. The structure design is simple and clear, which is convenient for hardware implementation and subsequent optimization. At the same time, this design has good scalability and can adapt to the requirements of multipliers with different bit-widths, making the design more flexible.
[0020] To make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the specific embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0022] Figure 1 It is a structural diagram of the FPGA multiplier optimization device of the present invention.
[0023] Figure 2 It is a schematic diagram of splicing multiple groups of small bit-width operation data of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] Regarding the foregoing and other technical contents, features, and effects of the present invention, they will be clearly presented in the following detailed description of a preferred embodiment in conjunction with the reference drawings. The directional terms mentioned in the following embodiments, such as: up, down, left, right, front, or back, etc., are only for reference to the directions of the attached drawings. Therefore, the directional terms used are for illustration and not for limiting the present invention.
[0025] Embodiment 1 Embodiment 1 provides an FPGA multiplier optimization method, which detects the actual effective bit widths of the data at two input ends, generates a bit width control signal, and based on the bit width control signal, selects a small bit width parallel calculation mode or a full bit width calculation mode; in the small bit width parallel calculation mode, multiple groups of input data with effective bits of small bit width are subjected to sign bit extension to generate parallel operation data including sign bits, and the generated parallel operation data is subjected to a shift operation and splicing and then sent into the multiplier for parallel operation; in the full bit width calculation mode, the data directly enters the multiplier for operation.
[0026] The following explains and illustrates an FPGA multiplier optimization method of the present invention in detail with reference to the accompanying drawings.
[0027] See Figure 1 , the FPGA multiplier optimization method in Embodiment 1 includes the following steps: Step 1: Detect the bit widths of the data data_ina and data_inb at two input ends. A leading zero detection circuit is used to detect the bit widths of the input data. The specific detection method is that the leading zero circuit captures the input data at the valid edge of the clock signal, scans bit by bit from the highest bit to the lowest bit of the data, counts the consecutive zeros until the first 1 is counted, outputs the leading zero detection result, and further determines the highest non-zero bit in the input data. The output leading zero detection result is used to represent the effective bit width of the input data.
[0028] Step 2: Generate a bit width control signal based on the leading zero detection result. Specifically, according to the leading zero detection results LZD_A and LZD_B, generate bit width control signals bit_width_a and bit_width_b, where bit_width_a represents the effective bit width of the input data data_ina, and bit_width_b represents the effective bit width of the input data data_inb. Then generate the bit width control signal through the following logic: bit_width_a = 18 - LZD_A; bit_width_b = 18 - LZD_B; The generation logic of the above bit width control signal is to subtract the leading zero detection result of each input data from the bit width 18 of an 18bit×18bit multiplier. This application uses an 18bit×18bit multiplier as the core of parallel multiplication calculation. The commonly used multiplier bit width in FPGA is 18. Of course, other bit width multipliers can also be used as the core of parallel multiplication calculation. The logical generation of the bit width control signal here can also be calculated according to the bit width of the selected multiplier.
[0029] Based on the above logic, the generated bit-width control signal can not only be used to represent the effective bit-width of the data, but also represent the shift amount required when the data is shifted (left shift).
[0030] Step S3: Based on the bit-width control signal, select the small-bit-width parallel calculation mode or the full-bit-width calculation mode. Specifically, a small-bit-width data merging threshold is preset in advance. When the generated bit-width control signal is less than or equal to the small-bit-width data merging threshold, the small-bit-width parallel calculation mode is started; if any one of the bit-width control signals of the two input-end data is greater than the small-bit-width data merging threshold, the full-bit-width calculation mode is started. In Embodiment 1, the set small-bit-width data merging threshold is 8. The above process judgment logic is as follows: When bit_width_a ≤ 8 and bit_width_b ≤ 8, start the small-bit-width parallel calculation mode; When bit_width_a > 8 or bit_width_b > 8, start the full-bit-width calculation mode.
[0031] Step S3-1: When the small-bit-width parallel calculation mode is selected, composite the small-bit-width valid-bit data.
[0032] Step S3-11: Perform bit-width detection and sign-bit extension processing on the input data, including performing sign-bit extension on the valid-bit data. The sign-bit extension means that the sign-bit extension occupies 1 bit. Specifically, see Figure 2 In, each group of input data after extension has one more bit than the original input data. Setting 1-bit sign-bit extension can support signed-bit operations, which is a basic operation for processing two's complement numbers, and can also prevent operation result overflow and sign errors.
[0033] The valid-bit data after sign-bit extension processing is subjected to a shift operation, including based on the actual bit-width of the detected input data and the shift amount of each group of data determined by the bit-width control signal, performing a shift operation on the valid-bit data after sign-bit extension processing. After the shift operation, multiple groups of small-bit-width operation data are obtained, and the multiple groups of small-bit-width operation data are combined together through a logical "OR" to align them to a specific segment of the multiplier. (Here, aligning to a specific field of the multiplier means that after the multiple groups of small-bit-width operation data are spliced, starting from the low bit, they are arranged bit by bit in sequence to the input end of the multiplier, and the high-bit vacant bits are filled with 0). The splicing of multiple groups of small-bit-width operation data is as Figure 2 shown, where Figure 2A schematic diagram of stitching small-bitwidth operation data with the same bitwidth is given. It should be noted that the stitching of small-bitwidth operation data with different bitwidths is similar to that of small-bitwidth operation data with the same bitwidth. The stitching of small-bitwidth operation data with different bitwidths can be directly referred to the stitching of small-bitwidth operation data with the same bitwidth, and this application will not repeat the explanation.
[0034] Step S3-12: To prevent overflow after stitching multiple groups of small-bitwidth operation data, dynamic control logic is used to control the stitching bitwidth. Specifically: Use dynamic control logic to calculate the total bitwidth of the stitched data. The calculation formula is as follows:
[0035] Among them, bit_width[i] is the effective bitwidth of the i-th segment of data, and +1 represents that the sign extension occupies 1 bit. According to the calculation result of the total bitwidth Total_Width, the number of stitched segments is dynamically controlled. Before each stitching, judge the relationship between the current remaining bitwidth and the small-bitwidth operation data to be stitched. If 18 - Total_Width < bit_width[i+1], stop stitching and send the stitched data to the subsequent module for calculation; if 18 - Total_Width ≥ bit_width[i+1], continue to shift and stitch. Through dynamic bitwidth limitation and stitching control logic, the problem of bitwidth overflow can be effectively avoided.
[0036] Step S3-2: When selecting the full-bitwidth calculation mode, the data is directly input into the multiplier.
[0037] Step S4: The data enters the multiplier for multiplication operation.
[0038] Step S4-1: The stitched multiple groups of small-bitwidth operation data are sent to the multiplier. After the stitching of multiple groups of small-bitwidth operation data is completed, starting from the low bit, they are arranged bit by bit to the input end of the multiplier in turn, and the high-bit vacant bits are filled with 0. Each group of small-bitwidth operation data is independently operated. For example Figure 2 In it, 3 groups of 4-bit small-bitwidth numbers are shifted and filled into data[14:10], [9:5], [4:0], and the highest vacant bit is filled with 0.
[0039] The Figure 1 , the input data dataa[data_ina1, data_ina2,...] and datab[data_inb1, data_inb2,...] after stitching processing are input into the parallel computing multiplication calculation module for calculation, and the calculation result result[result1, result2,...] is obtained.
[0040] After the multiplier performs parallel calculations, the output results are divided into multiple segments, and the operation result of each segment corresponds to the operation result of a group of small-bitwidth operation data; the shifter extracts the operation result of each segment, that is, the operation result of each group of small-bitwidth significant-bit data is obtained. Taking the multiplication with a 6-bit width as an example above, the extracted operation results are result1 = result[11:0], result2 = result[23:12], and result3 = result[35:24]. Each operation result represents the operation result of a group of small-bitwidth significant-bit data. By implementing the method in Example 1, the bitwidth multiplexing of the multiplier can be achieved, reducing the space vacancy of the multiplier and improving the usage efficiency.
[0041] Step S4-2: When the full-bitwidth calculation mode is selected, the data is directly input into the multiplier for calculation, and the calculation result is output.
[0042] Example 2 Example 2 provides an FPGA multiplier optimization device for implementing an FPGA multiplier optimization method described in Example 1. Refer to Figure 1 , the multiplier optimization device includes: an input bitwidth detection module that uses a leading zero detection circuit to detect the effective bitwidth of the input data data_ina and data_inb; generates a bitwidth control signal, and based on the bitwidth control signal, selects a small-bitwidth parallel calculation mode or a full-bitwidth calculation mode; specifically, selecting a calculation mode based on the bitwidth control signal includes comparing the effective bitwidth of the input data with the small-bitwidth data merging threshold to determine whether to start the small-bitwidth parallel calculation mode (that is, Figure 1 the bitwidth < 8-bit judgment module in Figure 1 ), and judging the relationship between the remaining bitwidth of the multiplier input data and the bitwidth of the multi-group of small-bitwidth operation data to be spliced to determine whether to continue data splicing (that is,
[0043] the 18 - Total_Width ≥ bit_width judgment module in ).
[0043] A bitwidth detection and sign-bit extension module, in the case of starting the small-bitwidth parallel calculation mode, performs bitwidth detection and sign-bit extension processing on the input data. The significant-bit data after sign-bit extension processing is subjected to a shift operation, including performing a shift operation on the significant-bit data after sign-bit extension processing based on the detected actual bitwidth of the input data and the preset shift amount of each group of data. After the shift operation, multiple groups of small-bitwidth operation data are obtained; A splicing judgment module judges the relationship between the remaining bitwidth of the multiplier input data and the bitwidth of the multi-group of small-bitwidth operation data to be spliced to determine whether to continue data splicing; a bitwidth splicing and multiplexing module, in the small-bitwidth parallel calculation mode, splices the small-bitwidth operation data to be spliced, and after the splicing of the multi-group of small-bitwidth operation data is completed, arranges them bit by bit in sequence from the low bit to the input end of the multiplier, and fills the vacant bits at the high bit with 0; A parallel multiplication calculation module performs parallel calculations on the concatenated small-bitwidth operation data input, outputs multiple segments of operation results, and each segment of the output result corresponds to the operation result of a group of small-bitwidth operation data; a shifter extracts each segment of the operation result to obtain the operation results of each group of small-bitwidth significant-bit data. When the full-bitwidth calculation mode is enabled, the data is directly input into the multiplier for operation to output the multiplication result.
[0044] Among them, Figure 1 sign_a and sign_b in represent the signed / unsigned situations of the two input data (data_ina and data_inb) respectively. When the value is the logical 1, it represents that the input data is a signed number, and when the value is the logical 0, it represents that it is an unsigned number. signa corresponds to data_ina, and signb corresponds to data_inb.
[0045] The above has introduced in detail a method and device for optimizing an FPGA multiplier provided by the present invention. Specific examples are used in this article to elaborate on the structure and working principle of the present invention. The description of the above embodiments is only used to help understand the method and core idea of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
Claims
1. An FPGA multiplier optimization method, characterized in that, Detect the actual effective bit widths of the data at two input terminals, generate a bit width control signal, and based on the bit width control signal, select a small bit width parallel calculation mode or a full bit width calculation mode; In the small bit width parallel calculation mode, perform sign extension on multiple groups of input data with a small bit width of valid bits to generate parallel operation data including sign bits, perform a shift operation on the generated parallel operation data, and after splicing, send it into a multiplier for parallel operation; In the full bit width calculation mode, the data directly enters the multiplier for operation.
2. The FPGA multiplier optimization method according to claim 1, wherein Use a leading zero circuit to judge the number of leading zeros in the input data, determine the highest non-zero bit in the input data, and then output a leading zero detection result to represent the effective bit width of the input data; Generate a bit width control signal according to the leading zero detection result; Select the corresponding calculation mode according to the bit width control signal and the set small bit width data merging threshold.
3. The method for optimizing an FPGA multiplier according to claim 2, wherein The leading zero circuit captures the input data at the effective edge of the clock signal, scans bit by bit from the highest bit to the lowest bit of the data, counts the consecutive zeros until the first 1 is counted, and outputs the leading zero detection result.
4. The method for optimizing an FPGA multiplier according to claim 2, wherein The bit width control signal is obtained by subtracting the leading zero detection result from the bit width of the multiplier where the data is input; If the bit width control signals of the data at both input terminals are less than or equal to the small bit width data merging threshold, start the small bit width parallel calculation mode; If any one of the bit width control signals of the data at the two input terminals is greater than the small bit width data merging threshold, start the full bit width calculation mode.
5. The optimization method of an FPGA multiplier according to claim 4, characterized in that The bit width of the multiplier is 18, and the small bit width data merging threshold is 8.
6. The optimization method of an FPGA multiplier according to claim 1, wherein When selecting the small bit width parallel calculation mode, compound the small bit width valid bit data, including: perform bit width detection and sign extension processing on the input data to obtain valid bit data including sign bits, and perform a shift operation on the valid bit data after sign extension processing, including based on the actually detected bit width of the input data and the preset shift amount for each group of data, perform a shift operation on the valid bit data after sign extension processing, after the shift operation, obtain multiple groups of small bit width operation data, and then splice and combine the multiple groups of small bit width operation data together through a logical "OR" to align to a specific segment of the multiplier; The spliced multiple groups of the small bit width operation data are sent into the multiplier, and each group of small bit width operation data performs independent operations.
7. An FPGA multiplier optimization method according to claim 6, characterized in that The bit width control signal determines the shift amount of the input data.
8. The optimization method of an FPGA multiplier according to claim 6, characterized in that, The output result of the multiplier is divided into multiple segments, and the output result of each segment corresponds to the operation result of a group of small bit width operation data; The shifter extracts the operation result of each segment to obtain the operation results of each group of small bit width valid bit data.
9. The FPGA multiplier optimization method according to claim 6, wherein Before splicing the small bit width operation data, use dynamic control logic to monitor the total bit width of the spliced multiple groups of small bit width operation data, including calculating the total bit width of the currently spliced multiple groups of small bit width operation data, and before each splicing, judge the relationship between the remaining bit width of the multiplier input data and the bit width of the small bit width operation data to be spliced. When the remaining bit width of the multiplier input data is greater than or equal to the bit width of the small bit width operation data to be spliced, continue the data splicing, otherwise stop the splicing and send the spliced data to the subsequent module for calculation.
10. A method for optimizing an FPGA multiplier according to claim 9, characterized in that, After multiple groups of small-bitwidth operation data are concatenated, starting from the low bit, they are arranged bit by bit in sequence to the input end of the multiplier, and the highest vacant bit is filled with 0.
11. An FPGA multiplier optimization device, characterized in that, For implementing an FPGA multiplier optimization method according to any one of claims 1-10, the multiplier optimization device includes: An input bitwidth detection module that uses a leading zero detection circuit to detect the bitwidth of the valid bits of the input data, generates a bitwidth control signal, and based on the bitwidth control signal, selects a small-bitwidth parallel calculation mode or a full-bitwidth calculation mode; A bitwidth detection and sign bit extension module that, in the small-bitwidth parallel calculation mode, performs bitwidth detection and sign bit extension processing on the input data. The valid bit data after sign bit extension processing is subjected to a shift operation, including a shift operation on the valid bit data after sign bit extension processing based on the detected actual bitwidth of the input data and the preset shift amount for each group of data. After the shift operation, multiple groups of small-bitwidth operation data are obtained; A concatenation judgment module that judges the relationship between the remaining bitwidth of the multiplier input data and the bitwidth of multiple groups of small-bitwidth operation data to be concatenated, and decides whether to continue data concatenation; A bitwidth concatenation and multiplexing module that, in the small-bitwidth parallel calculation mode, concatenates the small-bitwidth operation data to be concatenated, and after multiple groups of small-bitwidth operation data are concatenated, starting from the low bit, they are arranged bit by bit in sequence to the input end of the multiplier, and the vacant bits at the high bit are filled with 0; A parallel multiplication calculation module that performs parallel calculation on the concatenated small-bitwidth operation data of the input, outputs multiple segments of operation results, and each segment of output result corresponds to the operation result of a group of small-bitwidth operation data; a shifter extracts each segment of operation result to obtain the operation results of each group of small-bitwidth valid bit data; When the full-bitwidth calculation mode is enabled, the data is directly input into the multiplier for operation to output the multiplication result.
Citation Information
Patent Citations
Structured mixed bit-width multiplying method and structured mixed bit-width multiplying device
CN102591615A
Mixed bit width accelerator based on DSP and fusion calculation method
CN114239819A
Multiplying unit
CN114424161A
Floating point multiplier and multiplying method
JP1993165605A
System and Method for Implementing a Multiplication
US20130166616A1