Base-4 Booth Multiplier and Its Implementation Method, Arithmetic Circuit and Chip
Through the design of selecting controller and multiple carry save adder, parallel calculation of the base 4 Booth multiplier is realized, solving the problem of improving calculation speed and improving overall performance.
Patent Information
- Application Number
- CN202210402706.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-04-02
- Filing Date
- 2022-04-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-04-18
AI Technical Summary
How to implement a base 4 Booth multiplier to improve its overall performance, especially to improve computing speed during partial product generation and compression.
The selection controller is used to output different types of gate control signals, combine a multi-bit selector and a multi-carry-save adder, and calculate the carry output of the partial product in parallel, and use a carry adder with a carry chain to add data to sum.
By calculating the carry output of the partial product in parallel, the duration of the calculation process is shortened, and the calculation speed and overall performance are improved.
Smart Images

Figure CN114756203B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of circuits, and in particular, to a radix-4 Booth multiplier, an implementation method thereof, an arithmetic circuit, and a chip. Background Art
[0002] The radix-4 Booth multiplier is one of the commonly used circuits in digital circuit design. For example, the radix-4 Booth multiplier is often used in complex logic chips such as a central processing unit (CPU) and a graphics processing unit (GPU), and also often used in comprehensive design chips such as a microcontroller unit (MCU) and a field programmable gate array (FPGA). Generally, the multiplication operation can be divided into three steps: partial product generation, compressing the partial products into two-row vectors, and finally adding the two-row vectors. In partial product generation, the radix-4 Booth encoding is usually adopted, and the radix-4 Booth encoding can reduce the number of partial products of the multiplier by half.
[0003] Therefore, how to implement the radix-4 Booth multiplier and further improve the overall performance of the radix-4 Booth encoding multiplier has become a technical problem to be solved urgently. Summary of the Invention
[0004] In view of this, the embodiments of the present application provide a radix-4 Booth multiplier, an implementation method thereof, an arithmetic circuit, and a chip, so as to overcome all or part of the above technical defects.
[0005] In a first aspect, the embodiments of the present application provide a radix-4 Booth multiplier, which includes: a selection controller, configured to output any one of a zeroing gating control signal for indicating zeroing of the partial product, a positive 1 multiple gating control signal for indicating that the partial product is the multiplicand multiplied by positive 1, a negative 1 multiple gating control signal for indicating that the partial product is the multiplicand multiplied by negative 1, a positive 2 multiple gating control signal for indicating that the partial product is the multiplicand multiplied by positive 2, a negative 2 multiple gating control signal for indicating that the partial product is the multiplicand multiplied by negative 2, and a sign bit gating control signal for indicating that the partial product is the multiplicand multiplied by a negative multiple, according to the values of each bit of the multiplier; wherein, the multiplier and the multiplicand are N-bit binary numbers;
[0006] A multi-bit selector is configured to receive a zeroing strobe control signal indicating that a partial product is set to zero, and output a first selection result for making the partial product zero; receive a +1 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by +1, and output a second selection result for making the partial product the multiplicand multiplied by itself; receive a -1 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by -1, and output a third selection result for making the partial product the multiplicand multiplied by -1; receive a +2 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by +2, and output a fourth selection result for making the partial product the multiplicand multiplied by 2; receive a -2 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by -2, and output a fifth selection result for making the partial product the multiplicand multiplied by -2; and,
[0007] A multi-way carry-save adder is configured to determine the corresponding bit positions of N-bit partial products with base-4 Booth multiplication carry weights in the N / 2 groups at the 0th bit position to the (2N - 1)th bit position, and respectively compress the partial products at the 0th bit position to the (2N - 1)th bit position, and output two groups of 2N-bit data. The number of carry-save adders used for compression by the multi-way carry-save adder at the 0th bit position to the (2N - 1)th bit position is the sum of the number of partial products and the number of sign bits at the corresponding bit positions minus 2;
[0008] A carry adder with a carry chain is configured to add and sum the two groups of 2N-bit data.
[0009] In a second aspect, the present application provides a method for implementing a base-4 Booth multiplier, which includes: according to the values of each bit of the multiplier, output any one of a zeroing strobe control signal indicating that a partial product is set to zero, a +1 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by +1, a -1 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by -1, a +2 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by +2, a -2 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by -2, and a sign bit strobe control signal indicating that the partial product is the multiplicand multiplied by a negative multiple; wherein, the multiplier and the multiplicand are N-bit binary numbers;
[0010] Upon receiving a zero - select strobe control signal indicating that the partial product is set to zero, output a first selection result for making the partial product zero; upon receiving a +1 - multiple select strobe control signal indicating that the partial product is the multiplicand multiplied by +1, output a second selection result for making the partial product the multiplicand multiplied by itself; upon receiving a - 1 - multiple select strobe control signal indicating that the partial product is the multiplicand multiplied by - 1, output a third selection result for making the partial product the multiplicand multiplied by - 1; upon receiving a +2 - multiple select strobe control signal indicating that the partial product is the multiplicand multiplied by +2, output a fourth selection result for making the partial product the multiplicand multiplied by 2; upon receiving a - 2 - multiple select strobe control signal indicating that the partial product is the multiplicand multiplied by - 2, output a fifth selection result for making the partial product the multiplicand multiplied by - 2;
[0011] Determine the bit positions corresponding to the N - bit partial products with radix - 4 Booth multiplication carry weights in the N / 2 groups at bit positions from the 0th bit to the (2N - 1)th bit, and compress the partial products at bit positions from the 0th bit to the (2N - 1)th bit respectively, and output two groups of 2N - bit data. The number of carry - save adders used for compression by the multiplexed carry - save adder at bit positions from the 0th bit to the (2N - 1)th bit is the sum of the number of partial products at the corresponding bit positions and the number of sign bits minus 2;
[0012] A carry adder with a carry chain is used to add and sum the two groups of 2N - bit data.
[0013] In a third aspect, the present application provides an arithmetic circuit, and the arithmetic circuit includes a radix - 4 Booth multiplier according to any one of the embodiments of the first aspect.
[0014] In a fourth aspect, the present application provides a chip, and the chip includes an arithmetic circuit according to any one of the embodiments of the third aspect.
[0015] The embodiments of the present application provide a radix-4 Booth multiplier, its implementation method, arithmetic circuit, and chip. Due to the selection controller, any one of the output zero selection control signal, positive 1-fold selection control signal, negative 1-fold selection control signal, positive 2-fold selection control signal, negative 2-fold selection control signal, and sign bit selection control signal is selected; a multi-bit selector is used to output the first selection result, the second selection result, the third selection result, the fourth selection result, and the fifth selection result; and a carry-save adder is used to determine the corresponding bit positions of N / 2 groups of N-bit partial products with radix-4 Booth multiplication carry weights on the 0th bit to the (2N - 1)th bit, and compress the partial products on the 0th bit to the (2N - 1)th bit respectively, and output 2 groups of 2N-bit data; a carry adder with a carry chain is used to add and sum the 2 groups of 2N-bit data, thereby basically realizing the parallel calculation of the carry output of each bit position in the partial product for the summation operation, thereby shortening the duration of the entire calculation process and improving the calculation speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Some specific embodiments of the embodiments of the present application will be described in detail hereinafter with reference to the drawings in an exemplary but non-limiting manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:
[0017] Figure 1 is a schematic structural diagram of a selection controller for implementing a radix-4 Booth multiplier provided by an embodiment of the present application;
[0018] Figure 2 is a table of partial products of a radix-4 Booth encoding method in a selection controller for implementing a radix-4 Booth multiplier provided by an embodiment of the present application;
[0019] Figure 3 is a schematic structural diagram of a multi-bit selector for implementing a radix-4 Booth multiplier provided by an embodiment of the present application;
[0020] Figure 4 is a schematic principle diagram of a 32-bit adder for implementing a radix-4 Booth multiplier provided by an embodiment of the present application for summing 8 groups of 16-bit data;
[0021] Figure 5 is a schematic structural diagram of a carry-save adder in a 32-bit adder for implementing a radix-4 Booth multiplier provided by an embodiment of the present application;
[0022] Figure 6Schematic diagram of the principle of a radix-4 Booth multiplier provided by an embodiment of the present application for summing 16 groups of 16-bit data;
[0023] Figure 7 Schematic diagram of the structure of a carry-save adder in a 64-bit adder for implementing a radix-4 Booth multiplier provided by an embodiment of the present application;
[0024] Figure 8 Schematic diagram of the structure of a carry adder with a carry chain in a radix-4 Booth multiplier provided by an embodiment of the present application;
[0025] Figure 9 Schematic diagram of the circuit of the first preprocessing unit in the carry module of the carry adder with a carry chain in a radix-4 Booth multiplier provided by an embodiment of the present application;
[0026] Figure 10 Schematic diagram of the circuit of the second preprocessing unit in the carry module of the carry adder with a carry chain in a radix-4 Booth multiplier provided by an embodiment of the present application;
[0027] Figure 11 Schematic diagram of the carry chain of the carry adder with a carry chain in a radix-4 Booth multiplier provided by an embodiment of the present application;
[0028] Figure 12 Schematic flowchart of a method for implementing a radix-4 Booth multiplier provided by an embodiment of the present application. Detailed implementation manners
[0029] The following further describes the specific implementation of the embodiments of the present invention with reference to the accompanying drawings of the embodiments of the present invention.
[0030] Embodiment 1
[0031] Figure 1 Schematic diagram of the structure of a selection controller for implementing a radix-4 Booth multiplier provided by an embodiment of the present application. The selection controller for implementing a radix-4 Booth multiplier in this embodiment can be an independent hardware circuit structure, or can be a basic circuit unit structure of other devices such as a chip or a microprocessor. As Figure 1 shown, the selection controller for implementing a radix-4 Booth multiplier provided by an embodiment of the present application includes a clear selection control module 101, a positive 1-times selection control module 102, a negative 1-times selection control module 103, a positive 2-times selection control module 104, a negative 2-times selection control module 105, and a sign bit selection control module 106. The multiplier and the multiplicand are N-bit binary numbers.
[0032] Partial products of the radix-4 Booth encoding method are asFigure 2 As shown, every three adjacent bits of the multiplier B have eight combination ways. Different combination forms respectively represent one of the partial products being 0, ±A, ±2A, where A represents the multiplicand. Among them, the zero setting selection control module 101 is used to implement the gating control signal when the partial product is zero. The positive 1 - fold selection control module 102 is used to implement the gating control signal when the partial product is the multiplicand itself. The negative 1 - fold selection control module 103 is used to implement the gating control signal when the partial product is the negative number corresponding to the multiplicand itself. The positive 2 - fold selection control module 104 is used to implement the gating control signal when the partial product is the multiplicand multiplied by 2 times. The negative 2 - fold selection control module 10 is used to implement the gating control signal when the partial product is the multiplicand multiplied by - 2 times. The sign - bit selection control module 106 is used to implement the gating control signal when the partial product is negative.
[0033] Specifically, the zero setting selection control module 101 is used to output a zero - setting gating control signal for characterizing the zero setting of the partial product when the (i + 1) - th bit, the i - th bit, and the (i - 1) - th bit of the multiplier are all high - level or all low - level. Wherein, i is an integer greater than or equal to 0 and less than or equal to N - 1; the partial product is used to characterize the product of the (i + 1) - th bit, the i - th bit, and the (i - 1) - th bit of the multiplier and the multiplicand based on the radix - 4 Booth multiplication. For example, the zero - setting gating control signal for the multiplicand A and the multiplier B can be expressed as
[0034] Specifically, the positive 1 - fold selection control module is used to output a positive 1 - fold gating control signal for characterizing the partial product being the multiplicand multiplied by positive 1 when the (i + 1) - th bit, the i - th bit, and the (i - 1) - th bit of the multiplier are respectively low - level, high - level, and low - level, or respectively low - level, low - level, and high - level. For example, the positive 1 - fold gating control signal for the multiplicand A and the multiplier B can be expressed as Optionally, in an embodiment of the present application, for the convenience of the overall layout during circuit implementation, the positive 1 - fold gating control signal can sometimes also be expressed as
[0035] Specifically, the negative 1 - fold selection control module is used to output a negative 1 - fold gating control signal for characterizing the partial product being the multiplicand multiplied by negative 1 when the (i + 1) - th bit, the i - th bit, and the (i - 1) - th bit of the multiplier are respectively high - level, high - level, and low - level, or respectively high - level, low - level, and high - level. For example, the negative 1 - fold gating control signal for the multiplicand A and the multiplier B can be expressed as Optionally, in an embodiment of the present application, for the convenience of the overall layout during circuit implementation, the positive 1 - fold gating control signal can sometimes also be expressed as
[0036] Specifically, the positive 2-fold selection control module is used to output a positive 2-fold gating control signal indicating that the partial product is the multiplicand multiplied by positive 2 when the (i + 1)-th bit, the i-th bit, and the (i - 1)-th bit of the multiplier are low level, high level, and high level respectively. For example, the positive 2-fold gating control signal for the multiplicand A and the multiplier B can be expressed as Optionally, in an embodiment of the present application, for the convenience of the overall layout during circuit implementation, the positive 1-fold gating control signal can sometimes also be expressed as
[0037] Specifically, the negative 2-fold selection control module is used to output a negative 2-fold gating control signal indicating that the partial product is the multiplicand multiplied by negative 2 when the (i + 1)-th bit, the i-th bit, and the (i - 1)-th bit of the multiplier are high level, low level, and low level respectively. For example, the negative 2-fold gating control signal for the multiplicand A and the multiplier B can be expressed as Optionally, in an embodiment of the present application, for the convenience of the overall layout during circuit implementation, the negative 2-fold gating control signal can sometimes also be expressed as
[0038] Specifically, the sign bit selection control module is used to output a sign bit gating control signal indicating that the partial product is the multiplicand multiplied by a negative multiple when the (i + 1)-th bit, the i-th bit, and the (i - 1)-th bit of the multiplier are high level, high level, and low level respectively, or high level, low level, and high level respectively, or high level, low level, and low level respectively. For example, the sign bit selection control module for the multiplicand A and the multiplier B can be expressed as PROC_2A = SELB_M1A·SELB_M2A.
[0039] In the embodiments of the present application, since the selection controller includes a reset selection control module for outputting a reset strobe control signal, a positive 1 - fold selection control module for outputting a positive 1 - fold strobe control signal indicating that the partial product is the multiplicand multiplied by positive 1, a negative 1 - fold selection control module for outputting a negative 1 - fold strobe control signal indicating that the partial product is the multiplicand multiplied by negative 1, a positive 2 - fold selection control module for outputting a positive 2 - fold strobe control signal indicating that the partial product is the multiplicand multiplied by positive 2, a negative 2 - fold selection control module for outputting a negative 2 - fold strobe control signal indicating that the partial product is the multiplicand multiplied by negative 2, and a sign - bit selection control module for outputting a sign - bit strobe control signal indicating that the partial product is the multiplicand multiplied by a negative multiple, the output cases of the partial product when the (i + 1)-th bit, the i - th bit, and the (i - 1)-th bit of the multiplier take various values can be covered by the reset selection control module, the positive 1 - fold selection control module, the negative 1 - fold selection control module, the positive 2 - fold selection control module, and the negative 2 - fold selection control module. Thus, the direct strobe of the partial product of each multiplier and the multiplicand can be realized in parallel, without performing multiple step - by - step operations. Therefore, the duration of the entire calculation process can be shortened and the calculation speed can be improved.
[0040] Figure 3 FIG. 4 is a schematic structural diagram of a multi - bit selector for implementing a radix - 4 Booth multiplier provided by an embodiment of the present application. The multi - bit selector for implementing a radix - 4 Booth multiplier in this embodiment can be an independent hardware circuit structure or a basic circuit unit structure of other devices such as a chip or a microprocessor. As Figure 3As shown in the figure, the multi-bit selector for implementing a radix-4 Booth multiplier provided by an embodiment of the present application includes a zero setting module 301, a first reverse transmission selection gate module 302, a first forward transmission selection gate module 303, a second reverse transmission selection gate module 304, a second forward transmission selection gate module 305, and a first inverter 306. The zero setting module 301, the first reverse transmission selection gate module 302, the first forward transmission selection gate module 303, the second reverse transmission selection gate module 304, and the second forward transmission selection gate module 305 are connected to the first inverter 306 after being connected by the same line, and the output end of the inverter serves as the output end of the multi-bit selector. Among them, the zero setting module 301, the first reverse transmission selection gate module 302, the first forward transmission selection gate module 303, the second reverse transmission selection gate module 304, and the second forward transmission selection gate module 305 are connected by the same line, which can be understood as being connected to the input end of the first inverter 306 in a way of "wired AND". The so-called "wired AND" means connecting the outputs of multiple circuits with high impedance states by the same line to implement the "AND" logic. Only one of the multiple paths will be selected, and the selected path in the "wired AND" will be used for data transmission. Since the other paths are in a high impedance state when not selected, only the selected path will perform data transmission. This is a way to implement the selector. In this embodiment of the present application, the five output paths of the zero setting module 301, the first reverse transmission selection gate module 302, the first forward transmission selection gate module 303, the second reverse transmission selection gate module 304, and the second forward transmission selection gate module 305 are commonly wired ANDed to the input end of the first inverter, and with the cooperation of the control circuit, only one path is guaranteed to be selected among the five output paths, so as to implement the function of a five-to-one selector. Through this parallel connection of three-state logics, the area can be significantly reduced to improve the area utilization rate.
[0041] When the multiplier and the multiplicand are 16-bit binary numbers, the corresponding adder for implementing the radix-4 Booth multiplier is a 32-bit adder. Specifically, Figure 4 is a schematic diagram of the principle of a 32-bit adder for implementing a radix-4 Booth multiplier provided by an embodiment of the present application for summing 8 groups of 16-bit data. Among them, each piece of data is a partial product, which is used to represent the product of the (i + 1)-th bit, the i-th bit, and the (i - 1)-th bit of the multiplier and the multiplicand based on the radix-4 Booth multiplication; i is an integer greater than or equal to 0 and less than or equal to 15. Specifically, the multi-way carry-save adder is used to determine the corresponding bits of the 8 groups of 16-bit partial products with the carry weights of the radix-4 Booth multiplication on the 0-31st bit positions. Since the carry weights of the 8 groups of partial products are different, after ranking according to each carry weight, it forms as Figure 4The misaligned arrangement form shown. The multi-way carry-save adder compresses the partial products on the 0th to 31st bit positions respectively and outputs two groups of 32-bit data. The number of carry-save adders used for compression by the multi-way carry-save adder on the 0th to 31st bit positions is the sum of the number of partial products and the number of sign bits on the corresponding bit positions minus 2.
[0042] Figure 5 It is a schematic structural diagram of a multi-way carry-save adder in a 32-bit adder for implementing a radix-4 Booth multiplier provided by an embodiment of the present application. The multi-way carry-save adder is used to perform 8-2 data compression on eight groups of 16-bit data and output two groups of 32-bit data. The number of carry-save adders corresponding to each bit position of the 32-bit adder is the sum of the number of partial products and the number of sign bits on the corresponding bit position minus 2. For example, for the 14th to 18th bit positions, the number of carry-save adders is 7.
[0043] When the multiplier and the multiplicand are 32-bit binary numbers, the corresponding adder for implementing the radix-4 Booth multiplier is a 64-bit adder. Specifically, Figure 6 It is a schematic diagram of the principle of a 64-bit adder for implementing a radix-4 Booth multiplier provided by an embodiment of the present application for summing sixteen groups of 16-bit data. Each of the data is a partial product, which is used to represent the product of the (j + 1)th, jth, and (j - 1)th bit positions of the multiplier and the multiplicand based on radix-4 Booth multiplication; j is an integer greater than or equal to 0 and less than or equal to 31. Specifically, the multi-way carry-save adder is used to determine the corresponding bit positions of the sixteen groups of 32-bit partial products with radix-4 Booth multiplication carry weights on the 0th to 63rd bit positions. Since the carry weights of the sixteen groups of partial products are different, after ranking according to each carry weight, a misaligned arrangement form as shown in Figure 6 is formed. The multi-way carry-save adder compresses the partial products on the 0th to 63rd bit positions respectively and outputs two groups of 64-bit data. The number of carry-save adders used for compression by the multi-way carry-save adder on the 0th to 63rd bit positions is the sum of the number of partial products and the number of sign bits on the corresponding bit positions minus 2.
[0044] Figure 7Schematic diagram of a carry-save adder in a 64-bit adder for implementing a radix-4 Booth multiplier provided by an embodiment of the present application. The carry-save adder is used to perform 8-2 data compression on 16 groups of 16-bit data and output 2 groups of 64-bit data. The number of carry-save adders corresponding to each bit position of the 64-bit adder is the sum of the number of partial products and the number of sign bits at the corresponding bit position minus 2. For example, for the 14th to 18th bit positions, the number of carry-save adders is 7, for the 15th bit position, the number of carry-save adders is 6, for the 16th bit position, the number of carry-save adders is 8, for the 17th bit position, the number of carry-save adders is 7, and for the 18th bit position, the number of carry-save adders is 9.
[0045] Figure 8 Schematic diagram of a carry adder with a carry chain in a radix-4 Booth multiplier provided by an embodiment of the present application. The carry adder with a carry chain in this embodiment can be an independent hardware circuit structure or a basic circuit unit structure of other devices such as a chip or a microprocessor. As Figure 8 shown, when the multiplier and the multiplicand are 32-bit binary numbers, the carry adder with a carry chain in the radix-4 Booth multiplier provided by an embodiment of the present application includes N carry modules 10, where N is an integer less than or equal to 7. Each carry module corresponds to multiple bit positions in 2 groups of 64-bit data, where the 2 groups of 64-bit data are 16-bit binary numbers. For example, a carry module can correspond to 2, 3, or more bit positions in 2 groups of 64-bit data. It should be understood that the number of bit positions in 2 groups of 64-bit data corresponding to each carry module among the N carry modules 10 can be the same or different. Among them, the partial product is used to represent the product of the (j + 1)th, jth, and (j - 1)th bit positions of the multiplier and the multiplicand based on the radix-4 Booth multiplication; j is an integer greater than or equal to 0 and less than or equal to 31.
[0046] Among them, the nth carry module is connected to the (n - 1)th carry module to receive the inter-stage carry parameter output by the (n - 1)th carry module. Thus, based on the inter-stage carry parameter output by the (n - 1)th carry module, the inter-stage carry parameter of the nth carry module and the carry output of each bit position corresponding to the nth carry module are calculated. Among them, n is an integer greater than 1 and less than or equal to N.
[0047] Each carry module includes a preprocessing unit and multiple carry calculation units, and one carry calculation unit corresponds to one bit position of 2 groups of 64-bit data.
[0048] In this embodiment, the preprocessing unit included in the nth carry module is used to preprocess multiple bit positions in the corresponding two groups of 64-bit data.
[0049] Optionally, in an implementation manner of the present application, the preprocessing result includes: an in-group carry generation signal and an in-group carry propagation signal. The preprocessing unit included in the nth carry module is specifically configured to: perform an operation on each bit position in the corresponding two groups of 64-bit data to generate a carry generation signal and a carry propagation signal corresponding to each bit position; generate an in-group carry generation signal and an in-group carry propagation signal for each bit position based on the carry generation signal and the carry propagation signal of at least one corresponding bit position respectively.
[0050] Specifically, perform a logical AND operation on each bit position in the corresponding two groups of 64-bit data to generate a carry generation signal for each bit position, and the carry generation signal is the logical AND value operation result of the corresponding bit positions in the two groups of 64-bit data. Perform a logical OR operation on each bit position in the corresponding two groups of 64-bit data to generate a carry propagation signal for each bit position, and the carry propagation signal is the logical OR value operation result of the corresponding bit positions in the two groups of 64-bit data. For the convenience of the overall layout in circuit implementation, in the embodiments of the present application, sometimes the result of performing a logical NOT operation on the carry generation signal of each bit position is also referred to as the carry generation signal. Similarly, the result of performing a logical NOT operation on the carry propagation signal of each bit position is also referred to as the carry propagation signal.
[0051] After obtaining the carry generation signal and the carry propagation signal of each bit position corresponding to the nth carry module, the preprocessing unit included in the nth carry module can also perform a logical OR operation on the carry generation signals of adjacent multiple bit positions to generate an in-group carry generation signal, and the preprocessing unit included in the nth carry module can also perform a logical AND operation on the carry propagation signals of adjacent multiple bit positions to generate an in-group carry propagation signal. For the convenience of the overall layout in circuit implementation, in the embodiments of the present application, sometimes the result of performing a logical NOT operation on the in-group carry generation signal is also referred to as the in-group carry generation signal. Similarly, the result of performing a logical NOT operation on the in-group carry propagation signal is also referred to as the in-group carry propagation signal.
[0052] For example, for the i-th bit position in the first addend A and the second addend B, the carry generation signal G of the i-th bit position i = A i · B i , and the carry propagation signal P of the i-th bit position i = A i + B i . As described above, for the convenience of the overall layout in circuit implementation, the carry generation signal and the carry propagation signal of the i-th bit position are sometimes also respectively represented as or The carry generation signal G within the group from the j-th bit to the i-th bit i:j = G i + G i+1 + … + G i , and the carry propagation signal P within the group from the j-th bit to the i-th bit i:j = P i · P i+1 ·... · P i . As described above, for the sake of the overall layout during circuit implementation, the carry generation signal and the carry propagation signal within the group from the j-th bit to the i-th bit can sometimes also be expressed as and
[0053] In addition, G i:j = G i:k + G k-1:j , and, P i:j = P i:k · P k-1:j , where k is any bit located between the j-th bit and the i-th bit in the order from the lowest bit to the highest bit.
[0054] In this embodiment, the multiple carry calculation units included in the n-th carry module are used to perform operations according to the preprocessed result and the inter-stage carry parameter of the (n - 1)-th carry module, and generate the carry output of each bit corresponding to the n-th carry module and the inter-stage carry parameter of the n-th carry module.
[0055] Optionally, in an embodiment of the present application, each carry calculation unit included in the n-th carry module is specifically used to perform operations according to the carry generation signal and the carry propagation signal within the group corresponding to the corresponding bit and the inter-stage carry parameter of the (n - 1)-th carry module, and generate the carry output of the corresponding bit.
[0056] For the highest bit among the multiple bits corresponding to the n-th carry module, the carry calculation unit corresponding to this highest bit is further used to use the carry parameter obtained in the calculation of the carry output of the highest bit corresponding to the n-th carry module as the inter-stage carry parameter of the n-th carry module.
[0057] Among them, the carry parameter is an intermediate quantity obtained during the calculation of the carry output of each bit, and there is a preset relationship between the carry parameter and the carry output. The carry output of each bit can be obtained by performing an operation based on the carry parameter of this bit and the carry propagation signal of this bit. Specifically, the carry output of each bit is the logical AND operation result of the carry parameter of this bit and the carry propagation signal of this bit. For example, if the carry output of the i-th bit is C i, the carry propagation signal of the i-th bit is P i , the carry parameter of the i-th bit is Cp i , then the preset relationship is: C i = P i · Cp i .
[0058] If the highest bit among the multiple bits corresponding to the (n - 1)-th carry module is the (k - 1)-th bit, then when the multiple carry calculation units in the (n - 1)-th carry module calculate the carry output C k-1 of the (k - 1)-th bit, the carry parameter Cp k-1 is obtained as the inter-stage carry parameter of the (n - 1)-th stage. If the output result of the preprocessing unit of the n-th carry module includes the in-group carry generation signal G i:k and the in-group carry propagation signal P i-1:k , then the carry output of the i-th bit is C i = G i:k + P i:k-1 · Cp k-1 . In addition, since P i:k-1 · Cp k-1 = P i:k · P k-1 · Cp k-1 , therefore, C i = G i:k + P i:k · C k-1 also holds.
[0059] Since G i:k and P i:k can be obtained through the processing of the preprocessing unit, therefore, when the carry calculation unit corresponding to the i-th bit in the n-th carry module obtains the inter-stage carry parameter C k-1 of the (n - 1)-th carry module, the carry output or carry parameter of the i-th bit can be obtained through simple logical operations. In addition, since the preprocessing unit in the n-th carry module can preprocess the multiple bits corresponding to the n-th carry module to obtain the corresponding multiple in-group carry generation signals and in-group carry propagation signals, the multiple carry calculation units in the n-th carry module can calculate the carry output of each bit in parallel based on the corresponding in-group carry generation signals and in-group carry propagation signals, thereby improving the efficiency of carry calculation.
[0060] It should be understood that for the sake of the overall layout during circuit implementation, the carry parameter Cp k-1 and the carry output C k-1 are sometimes also represented as and
[0061] In the embodiment of the present application, since the preprocessing unit included in the nth carry module preprocesses multiple bits in two groups of corresponding 64-bit data, and the nth carry module includes multiple carry calculation units for performing operations according to the preprocessing result and the inter-stage carry parameter of the (n - 1)th carry module to generate the carry output of each bit corresponding to the nth carry module and the inter-stage carry parameter of the nth carry module. This enables each carry calculation unit in the nth carry module to directly utilize the preprocessing result and the inter-stage carry parameter output by the (n - 1)th carry module to calculate the carry output of each corresponding bit in parallel when obtaining the inter-stage carry parameter output by the (n - 1)th carry module. Thus, the carry output of each bit in the 16-bit binary data is basically calculated in parallel.
[0062] In addition, as Figure 8 shown, the multiple carry-save adders in the radix-4 Booth multiplier further include a summation module. The summation module is electrically connected to N carry modules for processing two groups of 64-bit data when the sign bit strobe control signal of the two groups of 64-bit data is valid. The processing includes: inverting the highest bit of all partial products of the multiplicand and the multiplier, adding 1 to the highest bit of the first partial product, and adding 1 bit number in front of the highest bit of all partial products, and the value of the bit number is 1; and for performing operations according to each bit in the processed two groups of 64-bit data and the corresponding carry output to obtain the corresponding summation result; where the sign bit strobe control signal is used to represent that the partial product is the multiplicand multiplied by a negative multiple.
[0063] For example, for the ith bit in the first addend A and the second addend B, the summation result of the ith bit can be obtained according to the following summation formula. The formula is:
[0064]
[0065] where C i-1 is the carry output of the (i - 1)th bit in the first addend A and the second addend A.
[0066] In this embodiment, since the carry output of each bit in the 16-bit binary data is basically calculated in parallel, the summation result of each bit in the 16-bit binary data can be basically calculated in parallel. Thus, the duration of the entire calculation process can be shortened and the calculation speed can be improved.
[0067] Optionally, in an embodiment of the present application, the number of bits in two groups of 64-bit data corresponding to the nth carry module is equal to or greater than the number of bits in two groups of 64-bit data corresponding to the (n - 1)th carry module.
[0068] Since the calculation of the carry output of each bit corresponding to the nth carry module depends on the inter-stage carry parameter of the (n - 1)th carry module, the carry operation time of each carry calculation unit in the nth carry module has a certain logical delay relative to the carry operation time of each carry calculation unit in the (n - 1)th carry module. By making the number of bits of the two groups of 64-bit data corresponding to the nth carry module equal to or greater than the number of bits of the two groups of 64-bit data corresponding to the (n - 1)th carry module, this logical delay can be fully utilized to calculate the in-group carry generation signal and the in-group carry propagation signal, avoiding the situation where the nth carry module waits for the inter-stage carry parameter of the (n - 1)th carry module during calculation, which is beneficial to further reducing the time consumed by the operation.
[0069] Optionally, in an embodiment of the present application, N is equal to 7. The 0th bit to the 3rd bit of the two groups of 64-bit data correspond to the 1st carry module, the 4th bit to the 7th bit of the two groups of 64-bit data correspond to the 2nd carry module, the 8th bit to the 15th bit of the two groups of 64-bit data correspond to the 3rd carry module, the 16th bit to the 31st bit of the two groups of 64-bit data correspond to the 4th carry module, the 32nd bit to the 48th bit of the two groups of 64-bit data correspond to the 5th carry module, the 49th bit to the 58th bit of the two groups of 64-bit data correspond to the 6th carry module, and the 50th bit to the 63rd bit of the two groups of 64-bit data correspond to the 7th carry module. Thus, the layout of the adder is relatively concentrated and the area is small, which is beneficial to the overall structured layout.
[0070] It can be understood that when the multiplier and the multiplicand are 16-bit binary numbers, the carry adder with a carry chain includes: M carry modules, each carry module corresponding to multiple bit positions of 2 sets of data of the 32 bits, where the m-th carry module is connected to the (m - 1)-th carry module for receiving the inter-stage carry parameter output by the (m - 1)-th carry module, the multiplier and the multiplicand are 16-bit binary numbers, M is an integer less than or equal to 5, and m is an integer greater than 1 and less than or equal to M; each carry module includes a preprocessing unit and multiple carry calculation units, and one carry calculation unit corresponds to one bit position of 2 sets of data of the 32 bits; wherein, the partial product is used to represent the product of the (i + 1)-th bit, the i-th bit, and the (i - 1)-th bit of the multiplier and the multiplicand based on the radix-4 Booth multiplication; i is an integer greater than or equal to 0 and less than or equal to 15. Specifically, when M is equal to 5, the 1st carry module corresponds to the 0th bit to the 3rd bit of 2 sets of data of the 32 bits, the 2nd carry module corresponds to the 4th bit to the 7th bit of 2 sets of data of the 32 bits, the 3rd carry module corresponds to the 8th bit to the 15th bit of 2 sets of data of the 32 bits, the 4th carry module corresponds to the 16th bit to the 23rd bit of 2 sets of data of the 32 bits, and the 5th carry module corresponds to the 24th bit to the 31st bit of 2 sets of data of the 32 bits.
[0071] It should be understood that in this embodiment, the number N of carry modules can be 2, 4, or more, and the specific bit positions corresponding to each carry module can be set as needed, and this embodiment does not limit this.
[0072] Based on the radix-4 Booth multiplier provided in Embodiment 1, further, this embodiment provides Figure 8 The structural schematic diagram of a carry module in the multiple carry-save adders in the radix-4 Booth multiplier shown. It should be understood that this carry module can be any one of the N carry modules in Embodiment 1. For the convenience of description, this carry module will be referred to as the n-th carry module hereinafter. In this embodiment, the preprocessing unit included in the n-th carry module includes at least one first preprocessing unit and at least one second preprocessing unit arranged alternately.
[0073] In this embodiment, the first preprocessing unit is used to perform an operation on the i-th bit and the (i - 1)-th bit in the corresponding 2 sets of data of 64 bits to generate a first preprocessing result, and the first preprocessing result indicates the logical OR operation result of the carry generation signals of the i-th bit and the (i - 1)-th bit, where i is an odd number.
[0074] Optionally, in a specific implementation manner of the present application, as Figure 9As shown, the first preprocessing unit includes: a first AND gate 201, a second AND gate 202, and a first NOR gate 203. The first input terminal and the second input terminal of the first AND gate 201 respectively receive the i-th bit. The output terminal of the first AND gate 201 is connected to the first input terminal of the first NOR gate 203. The first input terminal and the second input terminal of the second AND gate 202 respectively receive the (i - 1)-th bit. The output terminal of the second AND gate 202 is connected to the second input terminal of the first NOR gate 203. The output terminal of the first NOR gate 203 outputs the first preprocessing result. For example, if the first addend is A and the second addend is B, then the first preprocessing result is where G i and G i-1 are the carry generation signal of the i-th bit and the carry generation signal of the (i - 1)-th bit.
[0075] It should be understood that the first preprocessing unit can also be directly implemented by a structure such as an AND-OR-NOT gate. This embodiment does not make any limitations in this regard.
[0076] In this embodiment, the second preprocessing unit is used to perform an operation on the j-th bit and the (j - 1)-th bit in two groups of corresponding 64-bit data, and generate a second preprocessing result. The second preprocessing result indicates the logical AND operation result of the carry propagation signals of the j-th bit and the (j - 1)-th bit, where j is an even number.
[0077] Optionally, in a specific implementation manner of the present application, as Figure 10 shown, the second preprocessing unit includes: a first OR gate 301, a second OR gate 302, and a first NAND gate 303. The first input terminal and the second input terminal of the first OR gate 301 respectively receive the j-th bit. The output terminal of the first OR gate 301 is connected to the first input terminal of the first NAND gate. The first input terminal and the second input terminal of the second OR gate 302 respectively receive the (j - 1)-th bit. The output terminal of the second OR gate 302 is connected to the second input terminal of the first NAND gate 303. The output terminal of the first NAND gate 303 outputs the second preprocessing result. For example, if the first addend is A and the second addend is B, then the first preprocessing result is where P j and P j-1 are the carry propagation signal of the j-th bit and the carry propagation signal of the (j - 1)-th bit.
[0078] It should be understood that the second preprocessing unit can also be directly implemented by a structure such as an OR-AND-NOT gate. This embodiment does not make any limitations in this regard.
[0079] Correspondingly, the multiple carry calculation units included in the n-th carry module are used to obtain the carry output of the corresponding bit based on at least one first preprocessing result, at least one second preprocessing result, and the inter-stage carry parameter of the (n - 1)-th carry module.
[0080] Optionally, in an embodiment of the present application, the preprocessing unit included in the nth carry module further includes a third preprocessing unit and a fourth preprocessing unit. The third preprocessing unit performs operations on at least two adjacent ones of the first preprocessing results output by at least one first preprocessing unit and the second preprocessing results output by at least one second preprocessing unit to generate corresponding third preprocessing results and fourth preprocessing results. The third preprocessing result indicates the carry parameter between corresponding adjacent multiple bits, and the fourth preprocessing result indicates the logical AND operation result of the carry propagation signals between corresponding adjacent multiple bits. The multiple carry calculation units included in the nth carry module are configured to obtain the carry output of the corresponding bit position based on the third preprocessing result, the fourth preprocessing result, and the inter-stage carry parameter of the (n - 1)th carry module.
[0081] For example, the third preprocessing unit performs operations on the first preprocessing result and as well as the second preprocessing result to generate a carry parameter indicating between the 4th bit position and the 7th bit position The fourth preprocessing unit performs operations on the basis of the second preprocessing result and the second preprocessing result to generate the logical OR operation result of the carry generation signals indicating between the 3rd bit position and the 6th bit position, that is, a carry propagation signal within a group (that is, PAN_6_3). The corresponding carry calculation unit may obtain the carry output of the 7th bit position based on the third preprocessing result GON_7_4, the fourth preprocessing result PAN_6_3, and the inter-stage carry parameter of the (n - 1)th carry module.
[0082] Optionally, in an embodiment of the present application, the multiple carry calculation units included in the nth carry module include a first carry calculation unit corresponding to the ith bit position. The first carry calculation unit includes a third OR gate, a third AND gate, and a second NOR gate;
[0083] The first input terminal of the third OR gate is connected to the output terminal of the corresponding second preprocessing unit. The second input terminal of the third OR gate is connected to the inter-stage carry parameter output by the (n - 1)th carry module. The output terminal of the third OR gate is connected to the first input terminal of the third AND gate. The second input terminal of the third AND gate is connected to the output terminal of the corresponding first preprocessing unit. The output terminal of the third AND gate outputs the carry parameter of the ith bit;
[0084] The output terminal of the third AND gate is connected to the first input terminal of the second NOR gate. The second input terminal of the second NOR gate receives the carry propagation signal of the i-th bit. The output terminal of the second NOR gate is connected to the summation module to output the carry output of the i-th bit to the summation module.
[0085] Optionally, in an embodiment of the present application, the multiple carry calculation units further include a second carry calculation unit corresponding to the j-th bit. The second carry calculation unit includes a fourth OR gate and a second NAND gate.
[0086] The first input terminal of the fourth OR gate is connected to the output terminal of the corresponding second preprocessing unit. The second input terminal of the fourth OR gate is connected to the inter-stage carry parameter output by the (n - 1)-th carry module or the carry parameter of the (j - 1)-th bit. The output terminal of the fourth OR gate is connected to the first input terminal of the second NAND gate. The second input terminal of the second NAND gate receives the carry generation signal corresponding to the j-th bit. The output terminal of the second NAND gate is connected to the summation module to output the carry output of the j-th bit to the summation module.
[0087] In this embodiment, since the first preprocessing unit, the second preprocessing unit, the third preprocessing unit, and the fourth preprocessing unit in each carry module preprocess multiple bits in two groups of 64-bit data corresponding to each carry module, and each carry module includes multiple carry calculation units, when each carry module obtains the inter-stage carry parameter output by the previous carry module, the multiple carry calculation units in each carry module can directly use the preprocessing result and the inter-stage carry parameter output by the previous carry module to calculate the carry output of each corresponding bit in parallel. Thus, the carry output of each bit in 16-bit binary data is basically calculated in parallel.
[0088] As Figure 11 shown, the first carry module corresponds to the 0th bit to the 3rd bit of two groups of 64-bit data, the second carry module corresponds to the 4th bit to the 7th bit of two groups of 64-bit data, the third carry module corresponds to the 8th bit to the 15th bit of two groups of 64-bit data, the fourth carry module corresponds to the 16th bit to the 23rd bit of two groups of 64-bit data, and the fifth carry module corresponds to the 24th bit to the 31st bit of two groups of 64-bit data.
[0089] In addition, by regularly arranging the first preprocessing unit, the second preprocessing unit, the third preprocessing unit, the fourth preprocessing unit, the first carry calculation unit, and the second carry calculation unit, it is possible to improve the calculation speed of the radix-4 Booth multiplier while reducing the occupied area of the radix-4 Booth multiplier, and making the wiring more concentrated, which is beneficial to the overall structured layout.
[0090] It should be noted that Figure 11 This is only a specific example used to illustrate the carry chain of the multi-bit carry-save adder in the radix-4 Booth multiplier provided in this embodiment. According to actual needs, the number of carry modules can be 2, 4, or more, and the specific bit positions corresponding to each carry module can be set as needed. This embodiment does not make any limitations in this regard.
[0091] Embodiment III
[0092] Based on the radix-4 Booth multiplier provided in the above embodiment, an embodiment of the present application provides a method for implementing a radix-4 Booth multiplier. Figure 12 This is a flowchart of a method for implementing a radix-4 Booth multiplier provided in an embodiment of the present application. As Figure 12 shown, the method for implementing the radix-4 Booth multiplier includes:
[0093] S1201. According to the values of each bit of the multiplier, output any one of the zero-selection control signal used to represent the partial product being set to zero, the positive 1-fold selection control signal used to represent the partial product being the multiplicand multiplied by positive 1, the negative 1-fold selection control signal used to represent the partial product being the multiplicand multiplied by negative 1, the positive 2-fold selection control signal used to represent the partial product being the multiplicand multiplied by positive 2, the negative 2-fold selection control signal used to represent the partial product being the multiplicand multiplied by negative 2, and the sign-bit selection control signal used to represent the partial product being the multiplicand multiplied by a negative multiple;
[0094] S1202. When receiving the zero-selection control signal used to represent the partial product being set to zero, output the first selection result for making the partial product zero; when receiving the positive 1-fold selection control signal used to represent the partial product being the multiplicand multiplied by positive 1, output the second selection result for making the partial product the multiplicand multiplied by itself; when receiving the negative 1-fold selection control signal used to represent the partial product being the multiplicand multiplied by negative 1, output the third selection result for making the partial product the multiplicand multiplied by -1; when receiving the positive 2-fold selection control signal used to represent the partial product being the multiplicand multiplied by positive 2, output the fourth selection result for making the partial product the multiplicand multiplied by 2; when receiving the negative 2-fold selection control signal used to represent the partial product being the multiplicand multiplied by negative 2, output the fifth selection result for making the partial product the multiplicand multiplied by -2;
[0095] S1203. Determine the bit positions corresponding to the N-bit partial products with radix-4 Booth multiplication carry weights in the 0th bit to the (2N - 1)th bit of N / 2 groups, and compress the partial products in the 0th bit to the (2N - 1)th bit respectively, and output 2 groups of 2N-bit data;
[0096] S1204. A carry adder with a carry chain is used to add and sum the 2 groups of 2N-bit data.
[0097] The implementation method of the radix-4 Booth multiplier provided by the embodiments of the present application is used to implement the radix-4 Booth multiplier in the foregoing device embodiments, and has the beneficial effects of the corresponding device embodiments, which will not be elaborated here.
[0098] Embodiment 4
[0099] The embodiments of the present application provide an arithmetic circuit, which includes a radix-4 Booth multiplier provided according to any one of the foregoing Embodiment 1 and Embodiment 2. The principle and effect are similar, which will not be elaborated here.
[0100] Embodiment 5
[0101] The embodiments of the present application provide a chip, which includes the arithmetic circuit provided according to the foregoing Embodiment 4. The principle and effect are similar, which will not be elaborated here.
[0102] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0103] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A radix-4 Booth multiplier, characterized in that, Including: A selection controller, configured to output, according to the values of each bit of a multiplier, any one of a zeroing strobe control signal for indicating zeroing of a partial product, a positive 1 multiple strobe control signal for indicating that the partial product is the multiplicand multiplied by positive 1, a negative 1 multiple strobe control signal for indicating that the partial product is the multiplicand multiplied by negative 1, a positive 2 multiple strobe control signal for indicating that the partial product is the multiplicand multiplied by positive 2, a negative 2 multiple strobe control signal for indicating that the partial product is the multiplicand multiplied by negative 2, and a sign bit strobe control signal for indicating that the partial product is the multiplicand multiplied by a negative multiple; wherein, the multiplier and the multiplicand are N-bit binary numbers; A multi-bit selector, configured to output a first selection result for making the partial product zero when receiving the zeroing strobe control signal for indicating zeroing of the partial product; output a second selection result for making the partial product the multiplicand multiplied by itself when receiving the positive 1 multiple strobe control signal for indicating that the partial product is the multiplicand multiplied by positive 1; output a third selection result for making the partial product the multiplicand multiplied by -1 when receiving the negative 1 multiple strobe control signal for indicating that the partial product is the multiplicand multiplied by negative 1; output a fourth selection result for making the partial product the multiplicand multiplied by 2 when receiving the positive 2 multiple strobe control signal for indicating that the partial product is the multiplicand multiplied by positive 2; output a fifth selection result for making the partial product the multiplicand multiplied by -2 when receiving the negative 2 multiple strobe control signal for indicating that the partial product is the multiplicand multiplied by negative 2; And A multi-way carry-save adder, configured to determine the bit positions corresponding to the N-bit partial products with base 4 Booth multiplication carry weights in N / 2 groups on the 0th bit position to the (2N - 1)th bit position, and respectively compress the partial products on the 0th bit position to the (2N - 1)th bit position, and output two groups of 2N-bit data. The number of carry-save adders used for compression by the multi-way carry-save adder on the 0th bit position to the (2N - 1)th bit position is the sum of the number of partial products and the number of sign bits on the corresponding bit positions minus 2; A carry adder with a carry chain, configured to add and sum the two groups of 2N-bit data.
2. The radix-4 Booth multiplier according to claim 1, wherein N is 16; the carry adder with a carry chain includes: M carry modules, each carry module corresponding to multiple bit positions of two groups of 32-bit data. Among them, the mth carry module is connected to the (m - 1)th carry module for receiving the inter-stage carry parameter output by the (m - 1)th carry module. The multiplicand and the multiplier are 16-bit binary numbers. M is an integer less than or equal to 5, and m is an integer greater than 1 and less than or equal to M; each carry module includes a preprocessing unit and multiple carry calculation units, and one carry calculation unit corresponds to one bit position of the two groups of 32-bit data; Wherein, the partial product is used to represent the product of the (i + 1)th bit, the ith bit, and the (i - 1)th bit of the multiplier and the multiplicand based on base 4 Booth multiplication; i is an integer greater than or equal to 0 and less than or equal to 15.
3. The radix-4 Booth multiplier according to claim 2, wherein M is equal to 5. The first carry module corresponds to the 0th bit to the 3rd bit of two groups of 32-bit data. The second carry module corresponds to the 4th bit to the 7th bit of two groups of 32-bit data. The third carry module corresponds to the 8th bit to the 15th bit of two groups of 32-bit data. The fourth carry module corresponds to the 16th bit to the 23rd bit of two groups of 32-bit data. The fifth carry module corresponds to the 24th bit to the 31st bit of two groups of 32-bit data.
4. The radix-4 Booth multiplier according to claim 1, wherein N is 32. The carry adder with a carry chain includes: K carry modules, each carry module corresponding to multiple bits of two groups of 64-bit data. Among them, the kth carry module is connected to the (k - 1)th carry module to receive the inter-stage carry parameter output by the (k - 1)th carry module. The multiplicand and the multiplier are 32-bit binary numbers. K is an integer less than or equal to 7, and k is an integer greater than 1 and less than or equal to K. Each carry module includes a preprocessing unit and multiple carry calculation units. One carry calculation unit corresponds to one bit of the two groups of 64-bit data. Among them, the partial product is used to represent the product of the (j + 1)th bit, the jth bit, and the (j - 1)th bit of the multiplier and the multiplicand based on the radix-4 Booth multiplication; j is an integer greater than or equal to 0 and less than or equal to 31.
5. The radix-4 Booth multiplier according to claim 4, characterized in that, k is equal to 7. The first carry module corresponds to the 0th bit to the 3rd bit of two groups of 64-bit data. The second carry module corresponds to the 4th bit to the 7th bit of two groups of 64-bit data. The third carry module corresponds to the 8th bit to the 15th bit of two groups of 64-bit data. The fourth carry module corresponds to the 16th bit to the 31st bit of two groups of 64-bit data. The fifth carry module corresponds to the 32nd bit to the 48th bit of two groups of 64-bit data. The sixth carry module corresponds to the 49th bit to the 58th bit of two groups of 64-bit data. The seventh carry module corresponds to the 50th bit to the 63rd bit of two groups of 64-bit data.
6. The radix-4 Booth multiplier according to any one of claims 1-5, characterized in that, The selection controller includes: A clear selection control module, used to output a clear gating control signal for characterizing that the partial product is cleared when three adjacent bits of the multiplier from high to low are all high level or all low level; A positive 1-times selection control module, used to output a positive 1-times gating control signal for characterizing that the partial product is the multiplicand multiplied by positive 1 when three adjacent bits of the multiplier from high to low are respectively low level, high level, and low level, or respectively low level, low level, and high level; A negative 1-times selection control module, used to output a negative 1-times gating control signal for characterizing that the partial product is the multiplicand multiplied by negative 1 when three adjacent bits of the multiplier from high to low are respectively high level, high level, and low level, or respectively high level, low level, and high level; A positive 2-times selection control module, configured to output a positive 2-times gating control signal for indicating that the partial product is the multiplicand multiplied by positive 2 when three adjacent bits of the multiplier from high to low are low level, high level, and high level respectively; A negative 2-times selection control module, configured to output a negative 2-times gating control signal for indicating that the partial product is the multiplicand multiplied by negative 2 when three adjacent bits of the multiplier from high to low are high level, low level, and low level respectively; A sign bit selection control module, configured to output a sign bit gating control signal for indicating that the partial product is the multiplicand multiplied by a negative multiple when three adjacent bits of the multiplier from high to low are high level, high level, and low level respectively, or high level, low level, and high level respectively, or high level, low level, and low level respectively.
7. The radix-4 Booth multiplier according to claim 6, wherein The multi-bit selector includes: A zero setting module, configured to receive a zero setting gating control signal for indicating zero setting of the partial product, and output a first selection result for making the partial product zero; A first reverse transmission selection gate module, configured to receive a positive 1-times gating control signal for indicating that the partial product is the multiplicand multiplied by positive 1, and output a second selection result for making the partial product the multiplicand multiplied by itself; A first forward transmission selection gate module, configured to receive a negative 1-times gating control signal for indicating that the partial product is the multiplicand multiplied by negative 1, and output a third selection result for making the partial product the multiplicand multiplied by -1; A second reverse transmission selection gate module, configured to receive a positive 2-times gating control signal for indicating that the partial product is the multiplicand multiplied by positive 2, and output a fourth selection result for making the partial product the multiplicand multiplied by 2; A second forward transmission selection gate module, configured to receive a negative 2-times gating control signal for indicating that the partial product is the multiplicand multiplied by negative 2, and output a fifth selection result for making the partial product the multiplicand multiplied by -2; and A first inverter, the zero setting module, the first reverse transmission selection gate module, the first forward transmission selection gate module, the second reverse transmission selection gate module, and the second forward transmission selection gate module are connected through the same line and then connected to the first inverter, and an output end of the inverter serves as an output end of the multi-bit selector.
8. A method for implementing a radix-4 Booth multiplier, characterized in that, It includes: According to the values of each bit of the multiplier, output any one of a zero setting gating control signal for indicating zero setting of the partial product, a positive 1-times gating control signal for indicating that the partial product is the multiplicand multiplied by positive 1, a negative 1-times gating control signal for indicating that the partial product is the multiplicand multiplied by negative 1, a positive 2-times gating control signal for indicating that the partial product is the multiplicand multiplied by positive 2, a negative 2-times gating control signal for indicating that the partial product is the multiplicand multiplied by negative 2, and a sign bit gating control signal for indicating that the partial product is the multiplicand multiplied by a negative multiple; wherein, the multiplier and the multiplicand are N-bit binary numbers; Upon receiving a zeroing strobe control signal indicating that a partial product is to be zeroed, output a first selection result for making the partial product zero; upon receiving a +1 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by +1, output a second selection result for making the partial product the multiplicand multiplied by itself; upon receiving a -1 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by -1, output a third selection result for making the partial product the multiplicand multiplied by -1; upon receiving a +2 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by +2, output a fourth selection result for making the partial product the multiplicand multiplied by 2; upon receiving a -2 multiple strobe control signal indicating that the partial product is the multiplicand multiplied by -2, output a fifth selection result for making the partial product the multiplicand multiplied by -2; Determine the bit positions corresponding to the N-bit partial products with radix-4 Booth multiplication carry weights in the N / 2 groups at bit positions from the 0th bit to the (2N - 1)th bit, and compress the partial products at bit positions from the 0th bit to the (2N - 1)th bit respectively, and output two groups of 2N-bit data. The number of carry-save adders used for compression by the multi-path carry-save adder at bit positions from the 0th bit to the (2N - 1)th bit is the sum of the number of partial products at the corresponding bit positions and the number of sign bits minus 2; A carry adder with a carry chain is used to add and sum the two groups of 2N-bit data.
9. An operation circuit, characterized in that The arithmetic circuit includes a radix-4 Booth multiplier according to any one of claims 1 to 7.
10. A chip, characterized in that, The chip includes the arithmetic circuit according to claim 9.
Citation Information
Patent Citations
Method and device for performing operations involving multiplication of selectively partitioned binary inputs using booth encoding
US20040225705A1
Fast determination of carry inputs from lower order product for radix-8 odd / even multiplier array
US5729485A