Pipelined multiplier

The pipeline multiplier optimizes neuromorphic computing by employing a modified Booth Radix-8 algorithm and a three-layer 'Wallace tree' pipeline to parallelize intermediate value generation, addressing inefficiencies in existing technologies and enhancing speed, area, and energy efficiency.

WO2025183589A1PCT designated stage Publication Date: 2025-09-04OBSHCHESTVO S OGRANICHENNOI OTVETSTVENNOSTIU MOTIV NEIROMORFNYE TEKHNOLOGII
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/RU2024/050330
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-27
Filing Date
2024-12-25
Publication Date
2025-09-04

AI Technical Summary

Technical Problem

Existing neuromorphic computing technologies face inefficiencies in performing multiplication-addition operations due to sequential execution, leading to reduced speed and increased energy consumption, and lack of compactness and energy efficiency.

Method used

A pipeline multiplier using a modified Booth Radix-8 algorithm for parallel generation of intermediate values, a three-layer 'Wallace tree' convolution pipeline, and operation selection multiplexing to perform multiplication and multiplication-addition operations efficiently, optimizing clock signal operations for increased speed and reduced resource usage.

Benefits of technology

The solution achieves faster processing of multiplication and multiplication-addition operations, reducing device area and energy consumption while maintaining high performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure RU2024050330_04092025_PF_FP_ABST
    Figure RU2024050330_04092025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to the field of computing devices. A pipelined multiplier comprises an intermediate value generator configured to be capable of receiving first and second operands and generating intermediate multiplication results, a Wallace tree-type triple-layer pipeline for the convolution of the intermediate values, and a multiplexer for selecting a multiplication or combined multiplication and addition operation with a third operand. The units are configured so that each subsequent unit operates on the inverse edge of a clock signal in relation to the preceding unit. The intermediate value generator is based on a Booth algorithm which generates all the intermediate values in parallel in a half-cycle. The technical result of the invention is an increase in the speed of the device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]PIPELINE MULTIPLIER The invention relates to the field of computer devices based on specific computing models, can be used in designing neuromorphic processor devices, and can also be used in central and graphic processor devices and other areas. The pipeline multiplier can be part of the arithmetic coprocessor of the processor core of a neuromorphic VLSI (very large-scale integrated circuit) and is intended to perform a multiplication operation on two binary operands or a combined multiplication-addition operation (MAC operation) on three operands. The two main operands are integers in the range from -32768 to 32767, 16 bits long, presented in additional code. The third additional operand for the MAC operation and the operand of the result of the operation are a 32-bit integer, presented in additional code. A pipeline arithmetic multiplier is known (see Russian patent No. 2546072,publ. 2015, G06F 7 / 527). The device contains two blocks that have a cellular structure and are divided into columns and rows, a memory element, elements for controlling the recording of information and reading information from the memory. The first block is an input block with a number of columns equal to the sum of the digits of the multiplicand and the multiplier and with a number of rows one greater than the number of digits of the multiplicand. The second block has a number of columns and a number of rows equal to the number of digits of the multiplicand and the multiplier, respectively. In this case, cells of four types (OR, AND, half adder, adder) with registers are used. The main disadvantage of this solution is the lack of implementation of the addition operation for implementing the multiplication-addition operation in the multiplier pipeline. Another disadvantage is the systolic architecture of the multiplier, which is slower compared to the architecture based on the Booth algorithm. A device for performing multiplication-addition operations is known (see US20190012143, published in 2019), including a multiplier, an adder,a first result register and a second result register connected to the outputs of the multiplier and the adder, respectively. It includes four selection units. The first selection unit is configured to selectively supply to the multiplier, and in response to the first control signal, a first value from the first set of values. The second selection unit is configured to selectively supply to the multiplier, and in response to the second control signal, a second value from the second set of values. The third selection unit is configured to selectively supply to the adder, and in response to the third control signal, a third value from the third set of values. The fourth selection unit is configured to selectively supply to the adder, and in response to the fourth control signal, a fourth value from the fourth set of values. The main disadvantage of this solution is the sequential implementation of the MAC operation, that is, the multiplication of two operands,and subsequent summation of the result of the multiplication with the third operand. Such a solution leads to a difference in the speed of execution of the multiplication and multiplication-addition operations, which complicates the work of the processor's operation queue controller, and also reduces the overall speed of execution of the multiplication-addition operation. The closest in technical essence to the claimed device is a multiplication-addition device and its operating method (US Patent No. 7,730,118, published in 2006, IPC G06F 7 / 38), containing an arithmetic unit for selective implementation of one of the multiplication or multiplication-addition instructions, including a multiplier, a convolution pipeline,a result register and an accumulator circuit. The multiplier is configured to receive the first and second operands and is capable of generating multiplication terms. An addition circuit for receiving the multiplication terms from the multiplier and the ability to combine them to obtain the multiplication result. A result register for receiving the multiplication result from the adder. An operation selection multiplexer circuit is connected to receive the value stored in the result register and an accumulation control signal that determines whether the arithmetic block performs a multiplication or a MAC operation. The known device uses an intermediate value generator based on the Booth algorithm, a six-stage Wallace tree convolution pipeline for 17 intermediate values, and a multiplication / combined multiplication-addition operation selection multiplexer. A disadvantage of the prototype is the use of the standard Booth algorithm, which generates a larger number of intermediate values,which increases the size of the pipeline. For current tasks of neuromorphic computing, the version of the calculator with a bit depth of 32 bits does not allow implementing a more compact, productive and energy-efficient solution. The objectives of the invention are to increase the speed, reduce the area of ​​the device and energy consumption when performing a multiplication or combined multiplication-addition operation. The technical result is a decrease in the number of generated intermediate values, due to which the speed increases, the area and energy consumption decrease. The technical result is achieved in that the pipeline multiplier contains a generator of intermediate values, configured to receive the first and second operands, and capable of generating intermediate multiplication results, a pipeline of convolution of intermediate values ​​of the "Wallace tree" type for adding the obtained intermediate values ​​and obtaining the multiplication result,multiplexer for selecting the operation of multiplication or combined multiplication-addition with the third operand, the bit depth of the input operands of the multiplier is 16 bits, the bit depth of the result of the operation is 32 bits, and as a result, the bit depth of the third operand for the MAC operation is also 32 bits, which is necessary for the efficient use of the device resources, the blocks for the clock signal are designed so that each subsequent block operates on the inverse edge of the clock signal, relative to the previous one, the generator of intermediate values ​​is based on the modified Booth Radix-8 algorithm, which implements parallel generation of all variants of intermediate values ​​in half a clock cycle, the three-layer pipeline adds 6 intermediate values. The intermediate value convolution pipeline is a three-layer pipeline of the "Wales tree" type,in which the intermediate values ​​are added together in pairs to obtain the result. The generation of intermediate value variants from the multiplicand operand occurs in the intermediate value generator blocks. The intermediate value generator multiplexer consists of six parallel multiplexers. The convolution pipeline performs three stages of addition, each of which is processed by a separate group of adders. In parallel with the generation of intermediate values ​​and their loading into the intermediate value convolution pipeline, the operation selection multiplexer, based on the value of the operation selection flag, by means of switching through the multiplexer, writes either the value of the 32-bit operand in the case of performing a combined multiplication-addition operation, or "logical 0" in the case of performing a multiplication operation into a 32-bit register,which then passes the value to the intermediate value convolution pipeline. The "16-bit multiplier input operand width" feature affects the technical result as follows: due to a shorter operand, fewer intermediate values ​​are generated, which reduces the number of adders in the first layer of the intermediate value convolution pipeline, which in turn leads to a decrease in the number of layers in the intermediate value convolution pipeline, which leads to an increase in speed, a decrease in the area and power consumption of the device. The 32-bit operation execution result width affects the technical result, since a decrease in the output operand width reduces the width of adders, registers, multiplexers and other elements of the device, which in turn leads to a decrease in power consumption and a decrease in the area of ​​the device. Execution of blocks according to the clock signal in such a way that each subsequent block operates on the inverse edge of the clock signal,relative to the previous one, leads to an increase in the frequency of the blocks, due to which an increase in the device performance is achieved. Parallel generation of intermediate values ​​in half a clock cycle allows pipeline processing of input operands, which leads to an increase in the device performance. The implementation of a three-layer pipeline adding 6 intermediate values ​​allows optimizing the structure of the intermediate value convolution pipeline, which leads to an increase in performance, a decrease in the area and power consumption of the device. A decrease in the number of generated intermediate values ​​is achieved by generating intermediate values ​​based on several elements of the multiplier operand at once. In the case of the Radix-8 variation, the multiplier is divided into bit "quartets", on the basis of which the coefficients are determined by which the multiplier is multiplied to obtain a set of intermediate values. The use of the Radix-8 algorithm increases the size of the applied redundant coding of the multiplier,which leads to a decrease in the number of intermediate values. In turn, a decrease in the bit depth of the input operands from 32 (see prototype) to 16 bits also leads to a decrease in the number of intermediate values. The cumulative effect of using these solutions allows us to reduce the width and depth of the convolution pipeline, since the pipeline adds only 6 intermediate values, instead of 17 in the prototype, which in turn leads to the optimization of the declared characteristics. Fig. 1 shows the schematic diagram of the device, where block 1 is the intermediate value generator, block 2 is the operation selection multiplexer, block 3 is the intermediate value convolution pipeline, input 4 is a 16-bit operand-multiplicand, input 5 is a 16-bit operand-multiplier, input 6 is logical 0, input 7 is a 32-bit operand for the MAC operation, input 8 is the MAC operation selection flag, output 9 is the result of the device operation. Fig. 2 shows the intermediate value generator,where 11 is the block of the generator of intermediate values ​​with a coefficient of 4, 12 is the block of the generator of intermediate values ​​with a coefficient of 3, 13 is the block of the generator of intermediate values ​​with a coefficient of 2, 14 is the block of the generator of intermediate values ​​with a coefficient of 1, 15 is the block of the generator of intermediate values ​​with a coefficient of minus 1, 16 is the block of the generator of intermediate values ​​with a coefficient of minus 2, 17 is the block of the generator of intermediate values ​​with a coefficient of minus 3, 18 is the block of the generator of intermediate values ​​with a coefficient of minus 4, block 19 is the multiplexer of the generator of intermediate values. Fig. 3 shows the block of the generator of intermediate values ​​with a coefficient of 1, where block 141 is an 18-bit register. Fig. 4 shows the block of the generator of intermediate values ​​with a coefficient of 2, where block 131 is an 18-bit register. In Fig. 5 shows a block of the intermediate value generator with a coefficient of 3,where block 121 is an 18-bit full binary adder. Fig. 6 shows the block of the intermediate value generator with a coefficient of 4, where block 111 is an 18-bit register. Fig. 7 shows the block of the intermediate value generator with a coefficient of minus 1, where block 151 is an 18-bit logical inverter, block 152 is an 18-bit full binary adder. Fig. 8 shows the block of the intermediate value generator with a coefficient of minus 2, where block 161 is an 18-bit logical inverter, block 162 is an 18-bit full binary adder. Fig. 9 shows the block of the intermediate value generator with a coefficient of minus 3, where block 171 is an 18-bit logical inverter, block 172 is an 18-bit full binary adder. In Fig. 10 shows the block of the intermediate value generator with a coefficient of minus 4, where block 181 is an 18-bit logical inverter, block 182 is an 18-bit full binary adder. Fig. 11 shows the circuit diagram of the intermediate value generator multiplexer, where blocks 191, 192, 193, 194, 195,196 - multiplexers. Fig. 12 shows the operation selection multiplexer, where block 21 is the multiplexer, block 22 is the 32-bit register. Fig. 13 shows the intermediate value convolution pipeline, where blocks 31, 32, 33, 34, 35, 36 are 32-bit full binary adders. In Fig. 14, the table shows the variations of bit quartets and the coefficients corresponding to them. In Fig. 15, the timing diagram shows the general operating principle of the invention. The device consists of three functional blocks: an intermediate value generator (block 1, see Fig. 2), an operation selection multiplexer (block 2, see Fig. 12) and a convolution block (block 3, see Fig. 13). Generation of intermediate value variants from the multiplicand operand (input 4) occurs in the intermediate value generator blocks (blocks 11-18, see Fig. 3-10). The intermediate value generator block with the coefficient "4" (block 11) writes the value of the multiplicand operand (input 4) into an 18-bit register, with a left shift of 2,thereby generating and storing an 18-bit value (block 111), which is subsequently transmitted to the input of block 19 (see Fig. 6). The intermediate value generator block with the coefficient "3" (block 12) feeds the value of the operand-multiplicand (input 4), with a left shift by 1 and an extension of the most significant bit by 1, and the value of the operand-multiplicand (input 4) with an extension of the most significant bit by 2, into the 18-bit adder (121), where an 18-bit value is created and stored by addition, which is subsequently transmitted to the input of block 19 (see Fig. 5). The intermediate value generator block with the coefficient "2" (block 13) writes the value of the operand-multiplicand (input 4) to the 18-bit register with a left shift by 1 and an extension of the most significant bit by 1, thereby generating and storing an 18-bit value (block 131), which is subsequently transmitted to the input of block 19 (see Fig. 4). The intermediate value generator block with the coefficient "1" (block 14) writes the value of the operand-multiplicand (input 4) to the 18-bit register,with the most significant bit extended by 2, thereby generating and storing an 18-bit value (block 141), which is then transferred to the input of block 19 (see Fig. 3). The intermediate value generator block with the coefficient "minus 1" (block 15) feeds the operand-multiplicand value (input 4) with the most significant bit extended by 2, into the 18-bit inverter (block 151), obtaining a value equal to the negative multiplicand minus one, then feeding it to the adder (block 152), where 1 is added to it, obtaining and storing an 18-bit value equal to the negative multiplicand, which is then transferred to the input of block 19 (see Fig. 7). The minus 2 intermediate value generator block (block 16) feeds the operand-multiplicand value (input 4), shifted left by 1 and extended in the high-order bit by 1, into the 18-bit inverter (block 161), producing a value equal to twice the negative multiplicand minus one, then feeding it into the adder (block 162), where 1 is added to it,receiving and storing the value of the doubled negative multiplicand, which is then transferred to the input of block 19 (see Fig. 8). The block of the generator of intermediate values ​​with the coefficient "minus 3" (block 17) feeds the value of the operand-multiplicand (input 4) with a left shift by 2, into the 18-bit inverter (block 171), receiving the value of the quadrupled negative multiplicand minus one, then feeding it into the adder (block 182), where the value of the operand-multiplicand (input 4) is added to it with an extension of the most significant bit by 2, and 1, receiving and storing the value of the negative tripled multiplicand, which is then transferred to the input of block 19 (see Fig. 9). The minus 4 intermediate value generator block (block 18) feeds the operand multiplicand value (input 4), shifted left by 2, into the 18-bit inverter (block 181), producing the quadruple negative multiplicand value minus one, then feeding it into the adder (block 182), where 1 is added to it,receiving and storing the value of the quadruple negative multiplicand, which is subsequently transmitted to the input of block 19 (see Fig. 10). The intermediate value generator multiplexer (block 19, see Fig. 11) loads the intermediate values ​​obtained from blocks 11-18 and the logical 0 line (input 6), which is equal to the result of multiplying the multiplicand by the zero coefficient, corresponding to the quartets (see Fig. 14), obtained from the operand-multiplier (input 5) into the intermediate value convolution pipeline (block 3). The intermediate value generator multiplexer consists of 6 parallel multiplexers (blocks 191-196). When transmitting an intermediate value to the convolution tree, the multiplexers normalize them to 32-bit capacity. Thus, the multiplexer (block 191) expands the most significant bit of the transmitted intermediate value by 14 bits, the multiplexer (block 192) shifts the value to the left by 3 and expands the most significant bit of the transmitted intermediate value by 11 bits,the multiplexer (block 193) shifts the value to the left by 6 and extends the most significant bit of the transmitted intermediate value by 8 bits, the multiplexer (block 194) shifts the value to the left by 9 and extends the most significant bit of the transmitted intermediate value by 5 bits, the multiplexer (block 195) shifts the value to the left by 12 and extends the most significant bit of the transmitted intermediate value by 2 bits, the multiplexer (block 196) shifts the value to the left by 15 and discards the most significant bit of the intermediate value. In parallel with the generation of intermediate values ​​(block 2) and their loading into the intermediate value convolution pipeline (block 3), the operation selection multiplexer (block 2), based on the value of the operation selection flag (input 8), by means of switching through the multiplexer (block 21), writes either the value of the 32-bit operand (input 7) in the case of executing a combined multiplication-addition operation, or a logical 0 (input 6) in the case of executing a multiplication operation into a 32-bit register (block 22),which then passes the value to the intermediate value convolution pipeline (block 3). The intermediate value convolution pipeline (block 3) is a three-layer pipeline of the "Wales tree" type (see Fig. 13, where blocks 31-36 are 32-bit full binary adders), in which the intermediate values ​​are added together in pairs to obtain a result, which is then passed to output 9. The choice of the tree structure of the convolution block is associated with the best performance of this type of structure compared to analogs and the possibility of streaming data processing. To increase the speed of the device, the operation of the stages was optimized relative to the clock signal: each subsequent stage of the multiplier operates on the inverse edge of the clock signal, relative to the previous one. Thus, the speed of the operation was achieved,equal to two clock cycles. The general operating principle of the invention is shown in the timing diagram in Fig. 15. At the start of operation, two 16-bit operands are fed to the inputs of the intermediate value generator (block 1): the multiplicand (input 4) and the multiplier (input 5). Based on the multiplicand operand, all variants of intermediate values ​​used in the algorithm are generated, with coefficients from 4 to minus 4 (blocks 11-18). In parallel, based on the multiplier operand, bit quartets are generated, on the basis of which intermediate values ​​are fed to the intermediate value convolution pipeline (block 3), obtained by selecting in accordance with the value of the bit quartets and normalizing to a 32-bit width in the multiplexers (blocks 191-196) according to the ordinal number of the quartet. Then,in the intermediate value convolution pipeline (block 3), parallel pairwise arithmetic summation of intermediate values ​​begins. After a group of intermediate values ​​enters the intermediate value convolution pipeline, the intermediate value generator (block 1) can start generating new intermediate values ​​from a new pair of operands (input 4 and input 5). In parallel to the first stage of convolution, a third 32-bit operand (input 7) can be added to the intermediate value convolution pipeline via the operation selection multiplexer (block 2) in the case of executing a combined multiplication-addition operation, otherwise this operand is taken to be equal to zero (input 6), and the multiplication operation is executed, the operation selection flag (input 8) controls the operation selection. Regardless of the operation being executed, the convolution pipeline performs three stages of addition,each of which is processed by a separate group of adders (blocks 31-36). After the third stage of convolution, the result of multiplication or combined multiplication-addition is transmitted to the output of device 9. Example of executing a multiplication operation. Let the operand-multiplicand (input 4) be equal to 0000000011111010 (hereinafter referred to as A), the operand-multiplier (input 5) be equal to 0000110010100010 (hereinafter referred to as B) and the operation selection flag (input 8) be equal to 0. In this case, the value of the MAC operation operand (input 7) is not considered, since the device executes a multiplication operation, and therefore, the reference logical 0 (input 6) is written to the operation selection multiplexer register (block 22). At the leading edge of the clock signal (hereinafter referred to as CK) in the intermediate value generator (block 1), the intermediate value generation blocks calculate intermediate 18-bit values ​​from A. Thus, block 11 generates the value 4*A equal to 000000001111101000, block 12 generates the value 3*A equal to 000000001011101110,Block 13 generates the value of 2*A equal to 000000000111110100, block 14 generates the value of A equal to 000000000011111010, block 15 generates the value of minus A equal to 111111111100000110, block 16 generates the value of minus 2*A equal to 111111111000001100, block 17 generates the value of minus 3*A equal to 111111110100010010, block 18 generates the value of minus 4*A equal to 111111110000011000. The intermediate value of the coefficient 0 is not generated, but is multiplied from the reference 0 (input 6). In parallel to this, the intermediate value generator multiplexer (block 19) performs switching of the intermediate operation convolution pipeline with the generation blocks in accordance with B (input 5). For this, B (input 5) is divided into 6 bit groups with mutual overlap of 1 bit, and are fed to the control inputs of the intermediate value generator multiplexers (blocks 191-196), and, at the control input of the multiplexer (block 191) the bit group is shifted to the left by 1,and at the control input of the multiplexer (block 196) the sign bit is expanded, thereby obtaining the following set of input quartets: control input of block 191 = 0100, control input of block 192 = 1000, control input of block 193 = 0101, control input of block 194 = 1100, control input of block 195 = 0001, control input of block 196 = 0000. Each coefficient of intermediate values ​​corresponds to one or more variants of quartets of binary numbers (see Fig. 14). Also, the intermediate value generator multiplexers (blocks 191-196) convert the intermediate values ​​to 32-bit form by shifting the value to the left by the quartet number multiplied by 3, and filling the remaining high-order bits by extending the sign bit, or vice versa, by discarding the high-order bit of the intermediate value in the case of quartet 5. Thus,the intermediate value generator multiplexer passes the following set of intermediate value pairs to the intermediate value convolution pipeline (block 3) into the first layer adders: adder block 31 0000000000000000000000111110100 and 11111111111111111110000011000000, adder block 32 00000000000000001011101110000000 and 11111111111111100000110000000000, adder block 33 0000000000001111101000000000000 and 000000000000000000000000000000000. After, at the falling edge of the CK clock signal, the first layer of adders (blocks 31-33) in the intermediate value convolution pipeline begins to add in pairs the values ​​transferred from the intermediate value generator (block 1). Thus, at the output of the first layer, 3 intermediate 32-bit values ​​are formed, which are transferred to the adders of the second layer: adder block 34 11111111111111111110001010110100 and 1111111111111111001101001110000000, adder block 35 000000000000111110100000000000. After,at the next leading edge of the CK clock signal, the second layer of adders (blocks 34-35) in the intermediate value convolution pipeline begins to add the intermediate values ​​transmitted from the previous layer in pairs. Since the multiplication operation is executed, the value input 6, equal to 00000000000000000000000000000000000, is fed to the input of adder block 35 from the operation selection multiplexer (block 2). Thus, at the output of the first layer, 2 intermediate 32-bit values ​​are formed, which are fed to adder block 36: 11111111111111001011011000110100 and 000000000000011111010000000000000. In parallel to this, the intermediate value generator (block 1) and the operation selection multiplexer (block 2) can begin processing the next set of operands. Afterwards, at the falling edge of the CK clock signal, the last layer of adders in the intermediate value convolution pipeline (block 3) completes the multiplication operation by adding the last intermediate values, forming the result (output 9),equal to 000000000000011000101011000110100. Then, until the next positive edge of CK, the result will be held at the output of the intermediate value convolution pipeline. An example of executing a combined multiply-add operation (MAC operation). Let the multiplicand operand (input 4) be equal to 0000000011111010 (hereinafter referred to as A), the multiplier operand (input 5) be equal to 0000110010100010 (hereinafter referred to as B) and the operation selection flag (input 8) be equal to 1, the MAC operation operand (input 7) be equal to 00001010111100110000101011110011, and therefore, in the operation selection multiplexer register (block 22, )the value of the MAC operation operand is written (input 7). At the rising edge of the clock signal (hereinafter referred to as CK) in the intermediate value generator (block 1), the intermediate value generation blocks calculate intermediate 18-bit values ​​from A. Thus, block 11 generates the value of 4*A equal to 000000001111101000, block 12 generates the value of 3*A equal to 000000001011101110, block 13 generates the value of 2*A equal to 000000000111110100, block 14 generates the value of A equal to 000000000011111010, block 15 generates the value of minus A equal to 111111111100000110, block 16 generates the value of minus 2*A equal to 111111111000001100, block 17 generates the value of minus 3*A equal to 111111110100010010, block 18 generates the value of minus 4*A equal to 111111110000011000. The intermediate value of the coefficient 0 is not generated, but is multiplied from the reference 0 (input 6).In parallel to this, the intermediate value generator multiplexer (block 19) switches the intermediate value convolution pipeline with the generation blocks in accordance with B (input 5). For this purpose, B (input 5) is divided into 6 bit groups with mutual overlap of 1 bit, and fed to the control inputs of the intermediate value generator multiplexers (blocks 191-196), wherein, at the control input of the multiplexer (block 191), the bit group is shifted to the left by 1, and at the control input of the multiplexer (block 196), the sign bit is expanded, thereby obtaining the following set of input quartets: control input of block 191 = 0100, control input of block 192 = 1000, control input of block 193 = 0101, control input of block 194 = 1100, control input of block 195 = 0001, control input of block 196 = 0000. Each coefficient of the intermediate values ​​corresponds to one or more variants of quartets of binary numbers (see Fig. 14).Also, the intermediate value generator multiplexers (blocks 191-196) convert the intermediate values ​​to 32-bit form by shifting the value to the left by the quartet number multiplied by 3, and filling the remaining high-order bits by extending the sign bit, or vice versa, by discarding the high-order bit of the intermediate value in the case of quartet 5. Thus, the intermediate value generator multiplexer (block 19) passes the following set of intermediate value pairs to the intermediate value convolution pipeline (block 3) into the adders of the first layer: adder block 31 00000000000000000000000111110100 and 11111111111111111110000011000000, adder block 32 000000000000000001011101110000000 and 111111111111111000001100000000000, adder block 33 00000000000001111101000000000000 and 000000000000000000000000000000000.Afterwards, at the falling edge of the CK clock signal, the first layer of adders (blocks 31-33) in the intermediate value convolution pipeline begins to add in pairs the values ​​transmitted from the intermediate value generator (block 1). Thus, at the output of the first layer, 3 intermediate 32-bit values ​​are formed, which are transferred to the adders of the second layer: adder block 34 1111111111111111110001010110100 and 111111111111111001101001110000000, adder block 35 00000000000001111101000000000000. After, at the next leading edge of the CK clock signal, the second layer of adders (blocks 34-35) begins to add the intermediate values ​​transferred from the previous layer in pairs into the intermediate value convolution pipeline. Since the MAC operation is being executed, a value (input 7) equal to 00001010111100110000101011110011 is fed from the operation selection multiplexer (block 2) to the input of adder block 35.Thus, at the output of the first layer, 2 intermediate 32-bit values ​​are formed, which are fed to the adder block 36: 11111111111111001011011000110100 and 00001011000000101010101011110011. In parallel to this, the intermediate value generator (block 1) and the operation selection multiplexer (block 2) can begin processing the next set of operands. Then, on the falling edge of the CK clock signal, the last layer of adders in the intermediate value convolution pipeline (block 3) completes the multiplication operation by adding the last intermediate values, forming a result (output 9) equal to 000010101111111110110000100100111. Then, until the next positive edge of CK, the result will be held at the output of the intermediate value convolution pipeline.The use of the claimed invention allows for the rapid execution of a large number of multiplication or combined multiplication-addition operations in a pipeline mode of operation, i.e., each cycle of operation, receiving a new set of operand values ​​at the input and issuing a new result at the output, obtained on the basis of the old set of values.

Claims

CLAIMS 1. A pipeline multiplier comprising an intermediate value generator configured to receive first and second operands and capable of generating intermediate multiplication results, a Wallace tree-type intermediate value convolution pipeline for adding the received intermediate values ​​and obtaining a multiplication result, a multiplexer for selecting a multiplication operation or a combined multiplication-addition with a third operand, characterized in that the bit depth of the input operands of the multiplier is 16 bits, the bit depth of the result of the operation is 32 bits, the bit depth of the third operand for the MAC operation is 32 bits, the blocks for the clock signal are designed so that each subsequent block operates on the inverse edge of the clock signal, relative to the previous one, the intermediate value generator is based on a modified Booth Radix-8 algorithm, accelerated generation of intermediate values ​​with non-unit coefficients occurs in half a cycle,parallel generation of all intermediate value variants, a three-layer pipeline adds six intermediate values.

2. The pipelined multiplier of claim 1, characterized in that the intermediate value convolution pipeline is a three-layer Wallace tree pipeline, in which the intermediate values ​​are added pairwise to obtain the result.

3. The pipelined multiplier of claim 1, characterized in that the generation of intermediate value variants from the multiplicand operand occurs in the intermediate value generator blocks.

4. The pipelined multiplier of claim 1, characterized in that the intermediate value generator multiplexer consists of six parallel multiplexers.

5. The pipelined multiplier of claim 1, wherein the convolution pipeline carries out three stages of addition, each of which is processed by a separate group of adders.

6. The pipelined multiplier of claim 1, wherein, in parallel with the generation of intermediate values ​​and their loading into the intermediate value convolution pipeline, the operation selection multiplexer, based on the value of the operation selection flag, by means of switching through the multiplexer, writes either the value of the 32-bit operand in the case of executing a combined multiply-add operation, or a logical zero in the case of executing a multiplication operation into a 32-bit register, which subsequently transfers the value to the intermediate value convolution pipeline.

Citation Information

Patent Citations

  • Conveyor arithmetic multiplier

    RU2546072C1

  • Pipeline module multiplier

    RU2797164C1

  • Circuit for performing a multiply-and-accumulate operation

    US10437558B2

  • Multiply-accumulate unit and method of operation

    US7730118B2