Multiplier device
By dividing the multiplier into multiple data blocks and performing product operation, the existing multiplier is solved in terms of computational efficiency and power consumption, and efficient multiplication operation is realized.
Patent Information
- Application Number
- CN202111388728.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-22
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2041-11-22
AI Technical Summary
Existing multipliers have shortcomings in terms of computing efficiency and power consumption, especially in mobile Internet of Things devices, it is difficult to improve multiplication computing efficiency while ensuring low power consumption.
A multiplier device is designed, by dividing the multiplier into multiple target data blocks and performing product operations with the multiplier in sequence, using the control module, the product compression module and the calculation result generation module to reduce the calculation cycle and improve the calculation efficiency.
The multiplication operation cycle is reduced, the calculation speed is improved, and the circuit area is kept small while reducing power consumption, adapting to the automatic adjustment of the operation cycle of leading zeros of different multipliers.
Smart Images

Figure CN114063972B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of chip design and FPGA, and particularly relates to a multiplier device. Background Art
[0002] In recent years, with the rapid development of technologies such as big data, Internet of Things, cloud computing, and edge computing, the application scope of computer systems has become wider and wider, and people have put forward higher and higher requirements for the data processing capabilities of chips. Not only the pursuit of higher circuit performance, but also the need to reduce power consumption and circuit area as much as possible. Among them, the multiplier, as an important computing component, is widely used in chips such as microprocessors, general-purpose processors, and digital signal processors. Especially when the computing requirements of mobile Internet of Things devices are increasing day by day, the development of battery technology has been stagnant. Therefore, there is an urgent need for a high-energy-efficiency general-purpose fast multiplier to improve the multiplication operation efficiency while ensuring relatively low power consumption. Summary of the Invention
[0003] Based on this, in view of the above technical problems, it is necessary to provide a multiplier device that can reduce power consumption and improve the multiplication operation efficiency.
[0004] A multiplier device includes a control module, a product compression module, and an operation result generation module, wherein,
[0005] The control module is configured to divide the input multiplier into multiple target data blocks and output each target data block in sequence, wherein each target data block respectively includes data with a first preset number of bits;
[0006] The product compression module is connected to the control module, and the product compression module is configured to perform product operations on the multiple target data blocks received in sequence with the multiplicand respectively to correspondingly obtain multiple compressed data;
[0007] The operation result generation module is connected to the product compression module, and the operation result generation module is configured to generate an operation result according to the multiple compressed data.
[0008] In one embodiment, the first preset number of bits is four bits, and the control module includes:
[0009] A shift register, configured to store the input multiplier, output the low four bits of the data in the shift register as the target data block, and shift the stored data four bits to the right after each output of the target data block;
[0010] A leading zero detector, connected to the shift register, configured to perform leading zero detection on the data in the remaining bits of the shift register after outputting the target data block, and generate a control signal according to the detection result, where the control signal is used to control the shift register to continue outputting and terminate outputting the low four bits of the data in the shift register.
[0011] In one embodiment, a leading zero detector is configured to perform an exclusive OR operation on the data in the remaining bits of the shift register and the data of a second preset number of bits, then perform an OR operation on each bit of the result of the exclusive OR operation, and generate a control signal according to the operation result. The control signal includes a loop instruction and a reset instruction, where the remaining number of bits in the shift register is the same as the second preset number of bits;
[0012] The loop instruction is used to control the shift register to continue outputting the target data block, and fill zeros in the high four bits after shifting the stored data four bits to the right;
[0013] The reset instruction is used to control the shift register to terminate the output of the target data block.
[0014] In one embodiment, the product compression module includes:
[0015] A multiplicand register for storing the multiplicand;
[0016] An AND logic array, respectively connected to the multiplicand register and the control module, for performing an AND logic operation on the four-bit data of the target data block and the multiplicand respectively to obtain four partial products correspondingly;
[0017] A tree compression sub-module, respectively connected to the AND logic array and the operation result generation module, for generating a compressed data according to the four partial products corresponding to the same target data block.
[0018] In one embodiment, the AND logic array includes a plurality of AND logic units, and each AND logic unit is configured to respectively obtain the data on each digit of the current target data block, and perform an AND logic operation on the data on each digit and the multiplicand to generate four partial products.
[0019] In one embodiment, the tree compression sub-module includes:
[0020] A first-level operation unit, connected to the AND logic array, for obtaining an operation product according to two of the four partial products, and obtaining another first operation product according to the remaining two partial products;
[0021] A second-level operation unit, respectively connected to the first-level operation unit and the operation result generation module, for obtaining a compressed data according to the two first operation products and transmitting it to the operation result generation module.
[0022] In one embodiment, the first-level operation unit includes:
[0023] A first-bit extender, connected to the AND logic array, for respectively performing bit extension on the four partial products to obtain four first extended products;
[0024] The first adder is connected to the first bit expander and the second stage operation unit respectively, and is used for performing addition operation on the four first expanded products, obtaining two first operation products and transmitting them to the second stage operation unit.
[0025] In one embodiment, the second-level operation unit includes:
[0026] A second bit expander is connected to the first-stage operation unit and is used to expand the number of bits of the two first operational products to obtain two second expanded products;
[0027] The second adder is connected to the second bit expander and the operation result generation module respectively, and is used for performing an addition operation on the two second extended products, obtaining a compressed data and transmitting it to the operation result generation module.
[0028] In one embodiment, the control module further includes:
[0029] a cycle counter connected to the leading zero detector and configured to count cycles according to the control signal and obtain a cycle signal;
[0030] The operation result generation module includes an alignment expansion unit, a third adder, an operation result register and a two-way distributor; wherein,
[0031] The alignment expansion unit includes two input ends and an output end, the two input ends of the alignment expansion unit are respectively connected to the cycle counter and the product compression module, and the output end is connected to the third adder. The alignment expansion unit is used to align and bit-expand the compressed data according to the periodic signal, obtain the aligned expanded data, and output it to the third adder;
[0032] The third adder includes two input terminals and one output terminal, and the two input terminals of the third adder are respectively connected to the alignment extension unit and the operation result register, and is used to perform an addition operation on the alignment extension data and the previous operation result data stored in the operation result register, obtain a current operation result, and transmit it to the two-way distributor;
[0033] The two-way distributor includes two input terminals and two output terminals. The input terminals of the two-way distributor are respectively connected to the leading zero detector and the third adder. One output terminal of the two-way distributor is connected to the operation result register. The output terminal of the other two-way distributor is used to output the current operation result. The two-way distributor is used to output the current operation result according to the control signal and store the operation result in the result register as an addend in the next addition operation of the third adder.
[0034] Operation result register, used to store operation results.
[0035] In one embodiment, the alignment extension unit includes:
[0036] Multiple bit splicers, connected to the product compression module, where the multiple bit splicers are respectively used to splice preset data into the compressed data to obtain multiple bit-spliced data;
[0037] A multiplexer, respectively connected to the cycle counter, multiple bit splicers and the third adder, is used to select one of the bit-spliced data according to the cycle signal and transmit the bit-spliced data as the aligned extended data to the third adder.
[0038] The above multiplier device includes a control module, a product compression module and an operation result generation module. Among them, the control module is used to divide the input multiplier into multiple target data blocks and sequentially output each of the target data blocks, where each of the target data blocks respectively includes data with a first preset number of bits; the product compression module is connected to the control module, and the product compression module is used to perform product operations on the multiple target data blocks received sequentially with the multiplicand to correspondingly obtain multiple compressed data; the operation result generation module is connected to the product compression module, and the operation result generation module is used to generate an operation result according to the multiple compressed data. According to the operation method, the traditional serial multiplier has the advantages of simple structure and small area, but its characteristic is that the generation of partial products and the accumulation of partial products are completed in stages. One partial product is generated in each cycle, and then a single adder is cyclically reused for accumulation. That is to say, for an n-bit multiplier, it takes n cycles to obtain the final result. Therefore, its operation cycle is very long, which is not conducive to the high-speed implementation of multiplication calculations. The present invention provides a fast multiplier device, which divides the multiplier into multiple target data blocks, sequentially operates the multiple target data blocks with the multiplicand, thereby obtaining multiple compressed data, and then generates an operation result according to the multiple compressed data. The operation cycle is reduced, and the operation efficiency is further improved. Description of the Drawings
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0040] Figure 1 It is one of the structural schematic diagrams of the multiplier device in an embodiment;
[0041] Figure 2 It is the second structural schematic diagram of the multiplier device in an embodiment;
[0042] Figure 3 It is the third structural schematic diagram of the multiplier device in an embodiment;
[0043] Figure 4 The fourth structural schematic diagram of the multiplier device in an embodiment;
[0044] Figure 5 The fifth structural schematic diagram of the multiplier device in an embodiment;
[0045] Figure 6 The structural schematic diagram of the alignment and extension unit in an embodiment;
[0046] Figure 7 The sixth structural schematic diagram of the multiplier device in an embodiment;
[0047] Figure 8 The seventh structural schematic diagram of the multiplier device in an embodiment. Detailed implementation manners
[0048] For ease of understanding this application, the following will describe this application more comprehensively with reference to relevant attached drawings. Embodiments of this application are shown in the attached drawings. However, this application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, the purpose of providing these embodiments is to make the disclosure of this application more thorough and comprehensive.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used in the description of this application herein are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0050] It can be understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish one element from another.
[0051] Spatial relationship terms such as "under", "below", "lower", "beneath", "above", "upper", etc. can be used herein to describe the relationship between one element or feature shown in the figure and other elements or features. It should be understood that in addition to the orientation shown in the figure, spatial relationship terms also include different orientations of the device during use and operation. For example, if the device in the attached drawing is flipped, the element or feature described as "under other elements" or "beneath it" or "under it" will be oriented "above" other elements or features. Therefore, the exemplary terms "under" and "below" can include both the upper and lower orientations. In addition, the device can also include other orientations (such as rotating 90 degrees or other orientations), and the spatial description terms used herein are accordingly interpreted.
[0052] It should be noted that when an element is considered to be "connected" to another element, it can be directly connected to the other element or connected to the other element through an intermediate element. In addition, in the following embodiments, "connection", if there is transmission of electrical signals or data between the connected objects, should be understood as "electrical connection", "communication connection", etc.
[0053] As used herein, the singular forms "a", "an" and "the" may also include the plural forms unless the context clearly dictates otherwise. It should also be understood that the terms "comprises / comprising", "has / including", etc. specify the presence of the stated features, wholes, steps, operations, components, parts, or combinations thereof, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, components, parts, or combinations thereof.
[0054] In the description of this specification, the descriptions referring to terms such as "some embodiments", "other embodiments", "ideal embodiments", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example.
[0055] In one of the embodiments, as Figure 1 shown, a multiplier device 100 is provided. The multiplier device 100 includes a control module 110, a product compression module 120, and an operation result generation module 130. The control module 110 is configured to divide the input multiplier into a plurality of target data blocks and output each target data block in sequence, where each target data block respectively includes data of a first preset number of bits; the product compression module 120 is connected to the control module 110, and the product compression module 120 is configured to perform product operations on the plurality of target data blocks received in sequence with the multiplicand respectively to obtain a plurality of compressed data correspondingly; the operation result generation module 130 is connected to the product compression module 120, and the operation result generation module 130 is configured to generate an operation result according to the plurality of compressed data.
[0056] Specifically, in this embodiment, according to the bit width n of the input multiplier, the multiplier is virtually divided into a plurality of target data blocks according to the first preset number of bits, for example, n / 4 data blocks. For example, for a 16-bit wide multiplier, the multiplier can be non-overlappingly divided into four target data blocks according to four bits. For example, if the multiplier is "11110000 1100 0011", the four target data blocks are respectively "1111", "0000", "1100", and "0011". It can be understood that the division of the 16-bit wide multiplier into four target data blocks in this embodiment is only for illustration and does not limit the protection scope of the present application.
[0057] In this embodiment, the multiplier is divided into multiple target data blocks, and the multiple target data blocks are successively operated on with the multiplicand to obtain multiple compressed data, and then the operation result is generated according to the multiple compressed data. The number of target data blocks in this embodiment is the number of operation cycles in this embodiment. Compared with the traditional binary serial multiplier, the number of operation cycles is reduced and the operation efficiency is improved.
[0058] In one embodiment, the first preset number of digits is four digits. As Figure 2 shown, a multiplier device 100 is provided. The control module 110 in the multiplier device 100 includes: a shift register 111 and a leading zero detector 112. The shift register 111 is configured to store the input multiplier, output the low four-bit data in the shift register as the target data block, and shift the stored data four bits to the right after each output of the target data block; the leading zero detector 112 is connected to the shift register 111 and is configured to perform a leading zero detection on the data in the remaining bits of the shift register after the output of the target data block, and generate a control signal according to the detection result. The control signal is used to control the shift register 111 to continue outputting and terminate the output of the low four-bit data in the shift register 111.
[0059] Among them, the leading zero is a format for displaying the 0s in front of a number. For example, the first four bits "0000" in the data "0000 1010" are called leading zeros. When the leading zero detector 112 detects that all the data in the remaining bits of the shift register are 0, it indicates that the multiplier in the shift register has been completely output. At this time, the leading zero detector 112 can generate a control signal with a low level, and at the rising edge of the clock, control the shift register 111 to terminate the output of the low four-bit data in the shift register; when the leading zero detector 112 detects that not all the data in the remaining bits of the shift register are 0, it indicates that the multiplier in the shift register has not been completely output. At this time, the leading zero detector 112 can generate a control signal with a high level, and at the rising edge of the clock, control the shift register 111 to continue outputting the low four-bit data in the shift register 111, and at the same time, fill the high four bits in the shift register 111 with 0s. It can achieve the purpose of effectively controlling the output data and terminating the data.
[0060] In one embodiment, the leading zero detector 112 is configured to perform a bitwise exclusive OR operation on the data in the remaining bits of the shift register 111 and the data of the second preset number of bits, then perform an OR operation on each bit of the result of the bitwise exclusive OR operation, and generate a control signal according to the operation result. The control signal includes a loop instruction and a reset instruction, where the remaining number of bits in the shift register 111 is the same as the second preset number of bits; the loop instruction is used to control the shift register 111 to continue outputting the target data block, and fill the high four bits with zeros after shifting the stored data four bits to the right; the reset instruction is used to control the shift register 111 to terminate the output of the target data block.
[0061] Specifically, the fact that the remaining number of bits in the shift register 111 is the same as the second preset number of bits means that when the remaining number of bits in the shift register 111 is 12 bits, the second preset number of bits is also 12 bits of 0. Perform a bitwise exclusive OR operation on the data in these 12 remaining bits and these 12 bits of 0, and then perform an OR operation on each bit of the result of the bitwise exclusive OR operation to obtain the comparison result and generate a control signal.
[0062] Among them, the control signal is a 1-bit voltage signal. During a period, when the control signal is at a high level, it is used as a loop instruction to control the shift register 111 to continue outputting the target data block, and fill the high four bits with zeros after shifting the stored data four bits to the right; during a period, when the control signal is at a high level, it is used as a reset instruction to control the shift register 111 to terminate the output of the target data block.
[0063] In one embodiment, as Figure 3 shown, a multiplier device 100 is provided. The product compression module 120 in the multiplier device 100 includes: a multiplicand register 121, an AND logic array 122, and a tree compression sub-module 123. The multiplicand register 121 is used to store the multiplicand; the AND logic array 122 is respectively connected to the multiplicand register 121 and the control module 110, and is used to perform an AND logic operation on the four-bit data of the target data block and the multiplicand respectively to obtain four partial products correspondingly; the tree compression sub-module 123 is respectively connected to the AND logic array 122 and the operation result generation module 130, and is used to generate a compressed data according to the four partial products corresponding to the same target data block.
[0064] Specifically, different target data blocks output by the control module 110 sequentially enter the product compression module. After a target data block performs an AND logic operation with the multiplicand, multiple partial products will be generated. After the multiple partial products pass through the tree compression sub-module 123, a compressed data corresponding to a target data block will be generated. Similarly, other target data blocks will also generate corresponding compressed data through the product compression module 120 until the output of the multiplier in the control module 110 is completed.
[0065] In one embodiment, the AND logic array 122 includes a plurality of AND logic units. Each AND logic unit is used to separately obtain the data of each digit in the current target data block, and perform an AND logic operation on the data of each digit and the multiplicand to generate four partial products.
[0066] Specifically, if the current target data block includes four-bit data, correspondingly, the AND logic array 122 includes four AND logic units. In this embodiment, first, the four-bit data in the target data block is divided into four separate data according to its high and low positions, one corresponding to one data, m[0], m[1], m[2], and m[3], where m[0] is the highest bit and m[3] is the lowest bit. For example, if the four-bit data of the target data block is "1010", from right to left, m[0] is "0", m[1] is "1", m[2] is "0", and m[3] is "1". Then, the four separate data m[0], m[1], m[2], and m[3] are respectively input into the corresponding AND logic units, and an AND logic operation is performed with the multiplicand to generate four partial products.
[0067] In one embodiment, as Figure 4 shown, a multiplier device 100 is provided. The tree compression sub-module 123 in the multiplier device 100 includes: a first-stage operation unit 1231 and a second-stage operation unit 1232. The first-stage operation unit 1231 is connected to the AND logic array 122 and is used to obtain an operation product according to two of the four partial products and obtain another first operation product according to the remaining two partial products; the second-stage operation unit 1232 is respectively connected to the first-stage operation unit 1231 and the operation result generation module 130, and is used to obtain a compressed data according to the two first operation products and transmit it to the operation result generation module 130.
[0068] Compared with the traditional array multiplier, this embodiment reduces one level of addition operation, and only two levels of addition operation are required to compress the four partial products into a compressed data.
[0069] In one embodiment, continue to refer to Figure 4 , the first-stage operation unit 1231 includes: a first-bit extender 1231A and a first adder 1231B. The first-bit extender 1231A is connected to the AND logic array 122 and is used to respectively perform bit extension on the four partial products to obtain four first extended products; the first adder 1231B is respectively connected to the first-bit extender 1231A and the second-stage operation unit 1232, and is used to perform an addition operation on the four first extended products to obtain two first operation products and transmit them to the second-stage operation unit 1232.
[0070] Specifically, as described in the above embodiments, there are high and low bits in the data at positions m[0], m[1], m[2], and m[3] in the target data block. After performing AND logic operations with the multiplicand respectively, high and low bit extensions need to be performed on them respectively. Among them, if not for wiring symmetry, according to the principle of manual multiplication calculation, low bit extensions can be performed on m[0], m[1], m[2], and m[3]. In this embodiment, for wiring symmetry, high bit extension can be performed on the operation results of m[0] and m[1] with the multiplicand, and low bit extension can be performed on the operation results of m[2] and m[3] with the multiplicand. Among them, the number of bits of the first adder 1231B in this embodiment is equal to the number of bits of the first extended product. For example, if the partial products in this embodiment are extended by two bits in both high and low bits respectively, the first adder 1231B in this embodiment is an (n + 2)-bit adder. Among them, the high bit extension of the operation results of m[0] and m[1] with the multiplicand means padding the front of the operation results of m[0] and m[1] with the multiplicand with multiple zeros; for example, if the operation result of m[0] with the multiplicand is "1010", the result of the first extender 1231A with two-bit high bit extension is "001010"; the low bit extension of the operation results of m[2] and m[3] with the multiplicand means padding the back of the operation results of m[2] and m[3] with the multiplicand with multiple zeros. For example, if the operation result of m[2] with the multiplicand is "1010", the result of the first extender 1231A with two-bit high bit extension is "1010 00". After extension, four corresponding first extended products are obtained. Then, two of the first extended products are selected for addition operation. For example, the two first extended products obtained by respectively performing high and low bit extensions on the operation results of m[0] and m[2] with the multiplicand are selected for addition operation to obtain a first-level operation product, and the two first extended products obtained by respectively performing high and low bit extensions on the operation results of the remaining m[1] and m[3] with the multiplicand are selected for addition operation to obtain another first-level operation product.
[0071] In one of the embodiments, continue to refer to Figure 4 , the second-level operation unit 1232 includes: a second extender 1232A and a second adder 1232B. The second extender 1232A is connected to the first-level operation unit 1231 and is used to perform bit extension on the two first operation products respectively to obtain two second extended products; the second adder 1232B is respectively connected to the second extender 1232A and the operation result generation module 130, and is used to perform addition operation on the two second extended products to obtain a compressed data and transmit it to the operation result generation module 130.
[0072] Specifically, similar to the above embodiments, in this embodiment, high - and low - order extensions need to be performed on the first - level operation products respectively. Taking m[0], m[1], m[2], and m[3] in the above - mentioned embodiments as an example, in this embodiment, the operation results obtained by m[0] and m[2] and the multiplicand are subjected to a first - level operation to obtain a first - level operation result, which is then subjected to a high - order extension to obtain a second extended product. The operation results obtained by m[1] and m[3] and the multiplicand are subjected to a first - level operation to obtain another first - level operation result, which is then subjected to a low - order extension to obtain another second extended product. Among them, as described in the above - mentioned embodiments, if the first extender 1231A extends by two bits, then in this embodiment, the number of bits of the second extender 1232A also extends by two bits to meet the carry requirement. And the second adder 1232B in this embodiment is an (n + 4) - bit adder, which thus satisfies the addition operation of the second extender 1232A in this embodiment after respectively extending the two first - level operation results by two digits, obtaining compressed data of n + 4 bits, and there is no need to consider the carry problem anymore.
[0073] In one embodiment, as Figure 5 shown, a multiplier device 100 is provided. The control module 110 in the multiplier device 100 further includes: a cycle counter 113. The cycle counter 113 is connected to the leading - zero detector 112. The cycle counter 113 is used to perform cycle counting according to a control signal to obtain a cycle signal. The operation - result generation module 130 includes an alignment and extension unit 131, a third adder 132, an operation - result register 133, and a two - way distributor 134.
[0074] Among them, the alignment and extension unit 131 includes two input terminals and one output terminal. The two input terminals of the alignment and extension unit 131 are respectively connected to the cycle counter 113 and the product compression module 120, and the output terminal is connected to the third adder 132. The alignment and extension unit 131 is used to align and bit-extend the compressed data according to the cycle signal, obtain the aligned and extended data and output it to the third adder 132; the third adder 132 includes two input terminals and one output terminal. The two input terminals of the third adder 132 are respectively connected to the alignment and extension unit 131 and the operation result register 133, and are used to perform an addition operation on the aligned and extended data and the previous operation result data stored in the operation result register 133, obtain the current operation result and transmit it to the two-way distributor 134; the two-way distributor 134 includes two input terminals and two output terminals. The input terminals of the two-way distributor 134 are respectively connected to the leading zero detector 112 and the third adder 132. One output terminal of the two-way distributor 134 is connected to the operation result register 133, and the other output terminal of the two-way distributor 134 is used to output the current operation result. The two-way distributor 134 is used to output the current operation result according to the control signal and store the operation result in the result register as the addend in the next addition operation of the third adder 132; the operation result register 133 is used to store the operation result.
[0075] Specifically, in this embodiment, the control signal generated by the leading zero detector 112 can either directly control whether the shift register continues or terminates the output of the low four bits of data, or obtain the cycle signal by sending a signal to the cycle counter. For example, when the shift register receives the cycle signal "11", it means that the current cycle has reached the maximum cycle number. At this time, the target data block has been completely output. The leading zero detector 112 detects that all the remaining data in the shift register is 0, and then terminates the output of the low four bits of data. If it is detected that the current cycle number is not the maximum cycle number, the output of the low four bits of data continues.
[0076] Among them, when the cycle signal output by the alignment extension unit 131 is "00", that is, when the cycle number is 0, (n - 4) bits of 0 are filled in the high bits of the (n + 4)-bit compressed data generated by the product compression module 120 to form 2n-bit data. For example, when n is 16, 12 bits of 0 are filled in the high bits of the 20-bit compressed data generated by the product compression module 120 to form 32-bit data. When the cycle counter 113 is "10", that is, when the cycle number is 1, (n - 4) bits of 0 are filled in the high bits of the (n + 4)-bit compressed data generated by the product compression module 120, and then it is shifted left by 4 bits, and 4 bits of 0 data are filled in the low bits. For example, when n is 16 and the 20-bit compressed data is "1010 1010 1010 1010 1010", after filling (n - 4) bits of 0 in the high bits, shifting left by 4 bits, and filling 4 bits of 0 data in the low bits, it forms "0000 0000 1010 1010 1010 1010 0000". And so on, when the cycle number of the cycle counter 113 is k (0 < k < n / 4), (n - 4) bits of 0 are filled in the high bits of the (n + 4)-bit compressed data generated by the product compression module 120, and then it is shifted left by 4k bits, and 4k bits of 0 data are filled in the low bits. It should be noted that the cycle signal "00" can represent a cycle number of 0, the cycle signal "01" can represent a cycle number of 1, the cycle signal "10" can represent a cycle number of 2, and the cycle signal "11" can represent a cycle number of 3, and so on. There can also be other conversion methods, not limited to the examples provided in this embodiment.
[0077] Specifically, the leading zero detector 112 first performs an exclusive OR operation on a bit-by-bit basis between the high (n - 4) bits of the multiplier shift register 111 and (n - 4) bits of 0 data, and then performs an OR operation on the (n - 4)-bit result to generate a 1-bit control signal. For example, when n is 16, the leading zero detection is to perform an exclusive OR operation on a bit-by-bit basis between the high 12 bits and 12 bits of 0 data, and then perform an OR operation on the 12-bit result to generate a 1-bit control signal. If the control signal is "0", the multiplication operation result is output at the rising edge of the clock, and the loop is terminated. If the control signal is "1", at the rising edge of the clock, the 2n-bit operation result is written into the operation result register 133, and at the same time, the multiplier shift register 111 is shifted right by 4 bits, and 4 high bits are filled with 0. For example, if the multiplier is "1010 1010 1010 1010", the result of shifting the multiplier shift register 111 right by 4 bits and filling 4 high bits with 0 is "0000 1010 1010 1010". And the cycle number of the cycle counter 113 is incremented by 1.
[0078] Specifically, the multiplier in this embodiment is divided into multiple target data blocks. It can be understood that there are high and low bits among the target data blocks in the multiplier. The operation result after multiplying the high - order target data block by the multiplicand and the operation result after multiplying the low - order target data block by the multiplicand need to be aligned and extended separately when added together. Therefore, in this embodiment, the compressed data obtained by the operation of the target data block and the multiplicand in each cycle needs to be aligned and extended through the alignment and extension unit 131. In addition, the operation result generation module 130 in this embodiment processes multiple compressed data in cycles. That is, after aligning and extending the compressed data in the first cycle, the first aligned and extended data obtained is first subjected to an addition operation with 2n 0 bits in the result register 133 to obtain the first operation result which is still equal to the compressed data, and this first operation result is transmitted to the two - way distributor 134. At this time, if the two - way distributor 134 receives the loop instruction in the control signal of the leading - zero detector 112, that is, when the aforementioned control signal is "1", the first operation result is transmitted to the operation result register 133, and the next compressed data is aligned and extended to obtain the second aligned and extended data, which is then transmitted to the third adder 132. At the same time, the operation result register 133 transmits the first operation result to the third adder 132. The third adder 132 adds the first operation result and the second operation result to obtain the third operation result, which is then transmitted to the two - way distributor 134. The operation is repeated until the two - way distributor 134 receives the termination instruction in the control signal of the leading - zero detector 112, that is, when the aforementioned control signal is "0", the final 2n - bit operation result is output, the cycle counter 113 is reset to 0, and the result register 133 is also reset to 0.
[0079] In one embodiment, as Figure 6 shown, a structural schematic diagram of an alignment and extension unit is provided. The alignment and extension unit 131 includes: a plurality of bit concatenators 1311 and a multiplexer 1312. The plurality of bit concatenators 1311 are connected to the product compression module 120. The plurality of bit concatenators 1312 are respectively used to splice preset data into the compressed data to obtain a plurality of bit - spliced data. The multiplexer 1312 is respectively connected to the cycle counter 113, the plurality of bit concatenators 1311, and the third adder 132, and is used to select one of the bit - spliced data according to the cycle signal and transmit the bit - spliced data as the aligned and extended data to the third adder 132.
[0080] Specifically, as described above, when the compressed data obtained by multiplying different target data blocks with the multiplicand is finally added, alignment is required. The bit splicer in the alignment extension unit 131 in this embodiment can achieve splicing and zero-padding at different positions to achieve the required alignment extension at the positions where different target data blocks are located. Then, through the cycle signal corresponding to the current target data block, the corresponding accurate alignment extension data is obtained through the multiplexer 1312. For example, in this embodiment, to implement the operation of the target data block and the multiplicand for four cycles, four bit splicers and a four-to-one multiplexer are required.
[0081] In one embodiment, as Figure 7 shown, a multiplier device 100 is provided. The multiplier device 100 in this embodiment includes a shift register 111, a leading zero detector 112, a cycle counter 113, a multiplicand register 121, an AND logic array 122, a first bit extender 1231A, a first adder 1231B, a second bit extender 1232A, a second adder 1232B, an alignment extension unit 131, a third adder 132, an operation result register 133, and a two-way distributor 134. The connection manner of the devices in this embodiment is the same as that in the above embodiment, and will not be elaborated here.
[0082] Compared with the traditional binary serial multiplier and parallel multiplier, this embodiment fully combines the advantages of the serial multiplier and the parallel multiplier. It not only has a smaller circuit area and is easy to wire, but also significantly improves the multiplication calculation speed.
[0083] In one embodiment, as Figure 8As shown, a multiplier device 100 is provided based on a multiplicand and a multiplier with a bit width of 16 bits. This embodiment includes a control module 110, a product compression module 120, and an operation result generation module 130. The control module 110 includes a 16-bit shift register 111, a leading zero detector 112, and a cycle counter 113. The product compression module 120 includes a 16-bit multiplicand register 121, four AND logic units, four first expanders 1231A. Among them, the first expander 1231A is divided into two low-order expanders and two high-order expanders, as well as two 18-bit first adders 1231B and a 20-bit second adder 1232B. Among them, the first adder 1231B is divided into an 18-bit adder m and an 18-bit adder n, and both the first adder 1231B and the second adder 1232B are parallel fast adders, as well as two second-order expanders 1232A. Among them, the second-order expander 1232A is divided into a high-order expansion and a low-order expansion. The operation result generation module 130 includes an alignment expansion unit 131, a 32-bit third adder 132, an operation result register 133, and a two-way distributor 134. The alignment expansion unit 131 includes four bit splicers and a four-to-one multiplexer.
[0084] Specifically, first, initialize. At the rising edge of the clock, input a 16-bit multiplicand into the multiplicand register 121, input a 16-bit multiplier into the multiplier shift register 111, reset the 2-bit cycle counter 113 to 0, and reset the operation result register 133 to 0.
[0085] The multiplier shift register 111 inputs the lower four bits of data m[0], m[1], m[2], and m[3] into four AND logic units of the AND logic array 122 respectively, and performs AND operations with the 16-bit data of the multiplicand register 121 to generate four 16-bit partial products. Moreover, for the symmetry of the design, the lower four bits of the multiplier shift register m are divided into two groups: m[2] and m[0], m[3] and m[1]. Among them, m[2] and m[0] are used as a group of data to perform AND logic operations with the data in the multiplicand register 121 respectively, and output two partial products. The partial product generated by m[2] is extended by 2 bits of 0 at the lower position, and the partial product generated by m[0] is extended by 2 bits of 0 at the higher position, generating two 18-bit partial products and inputting them into the 18-bit adder m for addition operation. At the same time, m[3] and m[1] are used as a group of data to perform AND operations with the data in the multiplicand register 121 respectively, and output two partial products. The partial product generated by m[3] is extended by 2 bits of 0 at the lower position, and the partial product generated by m[1] is extended by 2 bits of 0 at the higher position, generating two 18-bit data and inputting them into the 18-bit adder n for addition operation. Among them, the output result (carry C, 18-bit sum) of the 18-bit adder m and 1 bit of 0 are extended at the higher position in the order of {1 bit of 0, C, 18-bit sum}. For example, if the two addends in the 18-bit adder m are "100000 0000 0000 0000" and "10 0000 0000 0000 0000" respectively, there will be a carry, and the output result is "C,10 0000 0000 0000 0000" in the order of {1 bit of 0, C, 18-bit}, and is extended at the higher position to "0100 00000000 0000 0000". The output result (carry C, 18-bit sum) of the other 18-bit adder n and 1 bit of 0 are extended at the lower position in the order of {C, 18-bit sum, 1 bit of 0}. Similarly, if the output result is "C,10 0000 0000 00000000" and is extended at the lower position in the order of {C, 18-bit sum, 1 bit of 0}, it becomes "100 0000 0000 0000 00000". Two 20-bit data generated by the two 18-bit adders are input into the 20-bit adder for addition operation and output 20-bit compressed data, which is transmitted to the operation result generation module 130. According to the properties of multiplication operations, multiplying a 16-bit multiplicand by a 4-bit multiplicand requires at most 20-bit data. Therefore, the result generated by the 20-bit adder has no overflow risk and there is no need to consider the carry.
[0086] In the operation result generation module 130, the alignment and extension unit 131 is used to bit-concatenate the 20-bit compressed data generated in the product compression module 120 and 12-bit 0 according to different combination orders, finally generating four 32-bit data, and outputting one of the data to the 32-bit adder for addition operation with the 32-bit data in the operation result register 133 according to the 2-bit control signal output by the cycle counter 113 through a 4-to-1 multiplexer. At the same time, two-way distributors 134 send the addition operation result to the operation result register 133 or directly output it according to the 1-bit control signal generated by the leading zero detector 112 in the control module 110.
[0087] Among them, the leading zero detector 112 first performs an exclusive OR operation on the high 12 bits of the multiplier shift register 111 and 12-bit 0 data bit by bit, and then performs an OR operation on the 12-bit result to generate a 1-bit control signal. If the control signal is "0", then when the rising edge of the clock arrives, the two-way distributors 134 output the operation result, and at the same time reset the operation result register 133 to 0, reset the cycle counter to 0, input new multiplicand and multiplier, and terminate the loop. If the control signal is "1", then when the rising edge of the clock arrives, the two-way distributors 134 input the 32-bit adder operation result into the operation result register 133, and at the same time shift the multiplier shift register 4 bits to the right, fill the high 4 bits with 0, and increment the cycle counter by 1.
[0088] Among them, the alignment and extension module is implemented by a combinational logic circuit. When the output of the cycle counter 113 is 0, it fills the high 12 bits of the 20-bit compressed data generated by the product compression module 120 with 0 to form 32-bit data and outputs it. When the cycle counter is 1, it fills the high 12 bits of the 20-bit data generated by the partial product generation and compression module with 0 and shifts it 4 bits to the left, and fills the low 4 bits with 0 data. And so on. When the cycle counter is k (0 < k < 4), it fills the high 12 bits of the 20-bit compressed data generated by the product compression module 120 with 0 and shifts it 4k bits to the left, and fills the low 4k bits with 0 data. According to the different numbers of leading zeros of the multiplier, let the number of leading zeros of the multiplier be L. When L > 11, the multiplier operation only requires 1 cycle. When 7 < L < 12, the multiplier operation requires 2 cycles. When 3 < L < 8, the multiplier operation requires 3 cycles. When L < 4, the multiplier operation requires 4 cycles.
[0089] According to the operation mode, the current traditional binary multipliers include serial multipliers and parallel multipliers. Among them, the traditional serial multiplier has the advantages of simple structure and small area. However, its characteristic is that the generation of partial products and the accumulation of partial products are completed in stages. One partial product is generated in each cycle, and then an adder is recycled for accumulation. That is to say, for an n-bit multiplier, it takes n cycles to obtain the final result. Therefore, its operation cycle is very long, which is not conducive to the high-speed implementation of multiplication calculations. The parallel multiplier, on the other hand, directly synchronously unfolds each step of the multiplication operation in the chip, and processes the generation of partial products and the compression of partial products in parallel, thereby greatly improving the speed of multiplication calculations. However, its circuit is very complex, wasting wiring resources and occupying a large area, and the power consumption is also very large. It is an idea of trading space for time. Currently, the mainstream parallel multipliers include array multipliers, Wallace tree multipliers, look-up table multipliers, etc.
[0090] As can be seen from the above, the traditional serial multiplier and parallel multiplier both have their respective advantages and disadvantages, and for different processed data, neither of these two multipliers can adaptively adjust the calculation process. For example, when calculating manually, if most of the high bits of the multiplier are zero, we can omit the operations on the high bits, thus saving the operation time. However, for the traditional binary multiplier, whether it is calculating a number with all bits being 0 or all bits being 1, it needs to complete all the calculation steps, lacking flexibility.
[0091] Aiming at the above defects of the prior art, the present invention provides a fast multiplier architecture, implementation method and working method. It fully combines the advantages of the serial multiplier and the parallel multiplier, not only has a smaller circuit area and is easy to wire, but also significantly improves the speed of multiplication calculations. At the same time, it can also automatically adjust the number of operation cycles according to the number of leading zeros of the multiplier, further improving the operation efficiency.
[0092] The technical features of the above embodiments can be combined arbitrarily. For the sake of brief description, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combinations of these technical features do not conflict, they should be considered as the scope described in this specification.
[0093] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A multiplier device, characterized in that, It includes a control module, a product compression module, and an operation result generation module. Among them, the control module is used to divide the input multiplier into multiple target data blocks and sequentially output each of the target data blocks. Each of the target data blocks respectively includes data with a first preset number of digits; the first preset number of digits is four digits; the product compression module is connected to the control module. The product compression module is used to perform product operations on the multiple target data blocks received sequentially with the multiplicand respectively to correspondingly obtain multiple compressed data; the operation result generation module is connected to the product compression module. The operation result generation module is used to generate an operation result according to the multiple compressed data; the control module includes: a shift register, which is used to store the input multiplier, output the low four bits of data in the shift register as the target data block, and shift the stored data four bits to the right after each output of the target data block; a leading zero detector, which is connected to the shift register. It is used to perform a leading zero detection on the data in the remaining bits of the shift register after outputting the target data block, and generate a control signal according to the detection result. The control signal is used to control the shift register to continue outputting and terminate the output of the low four bits of data in the shift register. The leading zero detector is used to perform an exclusive OR operation on the data in the remaining bits of the shift register and the data with a second preset number of digits first, then perform an OR operation on each bit of the result of the exclusive OR operation, and generate a control signal according to the operation result. The control signal includes a loop instruction and a reset instruction, where the number of remaining bits in the shift register is the same as the second preset number of digits; the loop instruction is used to control the shift register to continue outputting the target data block, and fill zeros in the high four bits after shifting the stored data four bits to the right; the reset instruction is used to control the shift register to terminate the output of the target data block.
2. The device according to claim 1, characterized in that, the product compression module includes: a multiplicand register, which is used to store the multiplicand; an AND logic array, which is respectively connected to the multiplicand register and the control module. It is used to perform AND logic operations on the four bits of data of the target data block and the multiplicand respectively to correspondingly obtain four partial products; a tree compression sub-module, which is respectively connected to the AND logic array and the operation result generation module. It is used to generate one compressed data according to the four partial products corresponding to the same target data block.
3. The device according to claim 2, characterized in that, The AND logic array includes multiple AND logic units. Each AND logic unit is used to respectively obtain the data on each digit in the current target data block, and perform an AND logic operation on the data on each digit and the multiplicand to generate four partial products.
4. The device according to claim 3, characterized in that, The tree compression sub-module includes: a first-level operation unit, which is connected to the AND logic array. It is used to obtain an operation product according to two of the four partial products, and obtain another first operation product according to the remaining two partial products; The second-level operation unit is respectively connected to the first-level operation unit and the operation result generation module, and is configured to obtain one piece of the compressed data according to two of the first operation products and transmit the compressed data to the operation result generation module.
5. The device according to claim 4, characterized in that, The first-level operation unit includes: A first-bit extender, connected to the AND logic array, and configured to respectively perform bit extension on the four partial products to obtain four first extended products; A first adder, respectively connected to the first-bit extender and the second-level operation unit, and configured to perform an addition operation on the four first extended products to obtain two of the first operation products and transmit the two first operation products to the second-level operation unit.
6. The device according to claim 4, wherein The second-level operation unit includes: A second-bit extender, connected to the first-level operation unit, and configured to respectively perform bit extension on the two first operation products to obtain two second extended products; A second adder, respectively connected to the second-bit extender and the operation result generation module, and configured to perform an addition operation on the two second extended products to obtain one piece of the compressed data and transmit the compressed data to the operation result generation module.
7. The device according to claim 1, characterized in that, The control module further includes: A cycle counter, connected to the leading zero detector, and configured to perform cycle counting according to the control signal to obtain a cycle signal; The operation result generation module includes an alignment and extension unit, a third adder, an operation result register, and a two-way distributor; wherein, The alignment and extension unit includes two input ends and one output end. The two input ends of the alignment and extension unit are respectively connected to the cycle counter and the product compression module, and the output end is connected to the third adder. The alignment and extension unit is configured to perform alignment and bit extension on the compressed data according to the cycle signal to obtain alignment and extended data and output the alignment and extended data to the third adder; The third adder includes two input ends and one output end. The two input ends of the third adder are respectively connected to the alignment and extension unit and the operation result register, and are configured to perform an addition operation on the alignment and extended data and the previous operation result data stored in the operation result register to obtain a current operation result and transmit the current operation result to the two-way distributor; The two-way distributor includes two input ends and two output ends. The input ends of the two-way distributor are respectively connected to the leading zero detector and the third adder. One of the output ends of the two-way distributor is connected to the operation result register, and the other output end of the two-way distributor is configured to output the current operation result. The two-way distributor is configured to output the current operation result according to the control signal and store the operation result in the result register as an addend for the third adder in the next addition operation; The operation result register is configured to store the operation result.
8. The device according to claim 7, characterized in that The alignment and extension unit includes: A plurality of bit splicers, connected to the product compression module, and the plurality of bit splicers are respectively configured to splice preset data into the compressed data to obtain a plurality of bit-spliced data; A multiple-choice selector, which is respectively connected to the cycle counter, multiple bit splicers, and the third adder, is configured to select one of the bit splicing data according to the cycle signal and transmit the bit splicing data as the aligned extended data to the third adder.
Citation Information
Patent Citations
Structured mixed bit-width multiplying method and structured mixed bit-width multiplying device
CN102591615A