Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

44 results about "Partial product" patented technology

Partial product. A product formed by multiplying the multiplicand by one digit of the multiplier when the multiplier has more than one digit. Partial products are used as intermediate steps in calculating larger products. For example, the product of 67 and 12 can be calculated as the sum of two partial products, 134 (67 X 2) + 670 (67 X 10), or 804.

High-speed fixed-point multiplication circuit

The invention discloses a high-speed fixed-point multiplication circuit. The multiplication circuit is mainly composed of a multiplier coding module, a partial product generation module, a partial product compression module and a traveling wave carry adder module. The multiplier coding module is composed of a radix-4-Booth coding algorithm and an opposite number generation module, and is used for carrying out three-bit block coding on input multiplication data and generating an opposite number of a multiplicand in advance. The partial product generation module generates a plurality of groups of partial products with symbol extension according to the multiplicand coded signal. And the partial product compression module adopts an improved Wallace compression structure to perform layered compression on the partial product. And the traveling wave carry adder module sums the two groups of results output by compression and outputs a multiplication result. According to the invention, multiplication accumulation series can be reduced, the switching times of invalid signals in the circuit can be reduced, and the longest delay path in an operation link can be shortened, so that fixed-point multiplication which is high in speed, low in power consumption and more favorable for a comprehensive tool in structure is realized.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Circuit for approximate floating point fusion dot product operation

The invention discloses a circuit for approximate floating point fusion dot product operation. The circuit comprises an extraction module used for extracting sign bits, mantissa bits and exponent bits from four single-precision floating point numbers; the symbol integration module is used for integrating symbol bits and correcting mantissa bits; the index comparison module is used for calculating two groups of dot product effective indexes according to the index bits and generating three control quantities; the first multiplication module is used for generating a first group of approximate partial products according to a correction mantissa digit corresponding to a first control quantity larger value; the second multiplication module is used for generating a second group of approximate partial products according to the correction mantissa digit corresponding to the smaller value of the first control quantity, and performing arithmetic displacement and adding symbol compensation bits according to the size of the second control quantity; the fusion compression module is used for performing fusion compression on the first group of approximate partial products, the second group of approximate partial products after arithmetic shift and the symbol compensation bits to obtain a sum sequence and a carry sequence; and the result output module is used for generating a dot product operation result in combination with the third control quantity.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Structured sparse matrix multiplier realized based on FPGA primitive

The invention discloses a structured sparse matrix multiplier realized based on FPGA primitives, and belongs to the technical field of FPGA hardware acceleration and deep learning computing architecture. The multiplier is composed of a register area and a multiplication and addition area, and the core design thought is that a circuit is built in a customized mode by directly calling bottom layer physical resources based on FPGA primitives; the register area stores a dense matrix B and supports parallel reading of elements by utilizing the characteristic that LUT in an SLICEM can be configured to be double SRL16E; the multiplication and addition area refers to a partial product generation unit based on the 4-Booth algorithm and an improved GPC (4: 2) compressor structure, and efficient generation and rapid accumulation of partial products are achieved. According to the method, redundant loss caused by high-level HDL logic synthesis is avoided through precise physical resource binding of primitives, fine wiring constraint and structured sparse data characteristic adaptation.
Owner:UNIV OF ELECTRONICS SCI & TECH OF CHINA

Multiplier circuit

Multiplier Circuit This description relates to a circuit configured to perform a multiplication operation between a first value and a second value, the first and second values ​​each comprising up to N pieces of bits respectively, the circuit comprising: - an NMUL number of multiplier subcircuits configured to generate first partial products, each first partial product corresponding to a multiplication between a piece from among the pieces of the first value and a piece from among the pieces of the second value; - an adding part configured to generate a first intermediate sum and first excess bit values, from the first partial products;and a parallel adder configured to generate a value corresponding to the product between the first and second values ​​from the first intermediate sum and the first integer bit generated by the first adder circuit and the second processing circuit. Figure for the abbreviation: Fig. 5;
Owner:COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES

Approximate multiplier based on FPGA

The invention provides an approximate multiplier based on a field programmable gate array (FPGA), which comprises an asymmetric partial product generation module, a signal processing module and a signal processing module, and is characterized in that the asymmetric partial product generation module is used for processing a received N-bit multiplier and a received N-bit multiplicand to generate an asymmetric approximate partial product matrix of which the line number is halved; the single-step compression module is used for compressing the asymmetric approximate partial product matrix into two rows at one time; and the final summing module is used for summing the two rows of compressed data and outputting a final product. Through the collaborative design of asymmetric partial product generation and single-step compression, high-energy-efficiency and low-delay approximate multiplication is realized on an FPGA platform, and the method is suitable for application scenes such as resource-limited edge calculation and the like.
Owner:HAINAN UNIV

A vector floating-point multiplier-accumulator suitable for floating-point operations of various precisions

ActiveCN116521124BFloating pointLogisim
This invention discloses a vector floating-point multiply-adder suitable for floating-point operations of various precisions, comprising a first operation module, a second operation module, a third operation module, and a fourth operation module. The first operation module includes a partial product generation module, a Wallace network, a first inversion module, an exponent alignment module, a mantissa compound right shifter, a sticky logic module, and an exception pre-judgment module. The second operation module includes a 3:2 CSA adder, a CPA adder, an increment circuit, a GRS logic module, and a sign pre-judgment module. The third operation module includes a second inversion module, a leading zero detection module, a trailing zero detection module, a normalized compound left shifter module, a normalization correction module, a rounding preprocessing module, a fast GRS solver module, and an exponent adjustment module. The fourth operation module includes a mantissa increment logic module, an exponent increment logic module, a sign judgment module, an exception judgment module, and a control logic output module. This invention balances computational speed and chip area, enabling parallel execution of floating-point multiply-add operations of various precisions.
Owner:SOUTH CHINA UNIV OF TECH

Floating-point number multiplier, near memory computing circuit, high-bandwidth magnetic computing chip system and electronic equipment

The invention provides a floating-point number multiplier, a near memory computing circuit, a high-bandwidth magnetic computing chip system and electronic equipment, and relates to the technical field of circuits, and the floating-point number multiplier comprises a compression circuit and a partial product generation circuit; the floating-point number multiplier is used for realizing mantissa multiplication of a known floating-point number of a multiplier; the partial product generation circuit comprises: a one-out-of-four selector; the compression circuit comprises a compressor and a summator. The one-out-of-four selector is used for generating a partial product of a known multiplier and a multiplicand after NR4SD + coding; the compressor is used for carrying out parallel compression on the partial product output by the one-out-of-four selector; and the adder is used for accumulating the parallel compression results output by the compressor and outputting an accumulation result to finish mantissa multiplication of the known multiplier and the multiplicand. The storage efficiency can be improved, and the area and power consumption of a hardware circuit are reduced.
Owner:ICY TECHNOLOGY (BEIJING) CO LTD

Processing circuit architecture supporting multi-calculation precision dynamic switching

The invention relates to the technical field of integrated circuit design, in particular to a processing circuit architecture supporting multi-calculation precision dynamic switching, which comprises a control port module, a partial product generation module, a symbol compression module, an addition compression tree module and a final adder module. And single-cycle dynamic switching of various precisions is realized. The partial product generation module adopts a Booth coding algorithm to split an operation vector, and cooperates with boundary symbol selection logic to solve the problem of symbol expansion and auxiliary bit overlapping; the symbol compression module compresses the redundant extension bits through a preset coding mode; the addition compression tree module is formed by cascading multiple stages of compressors and inserting carry blocking logic; the final adder module is composed of a plurality of carry lookahead adders, and the output bit width is dynamically controlled through blocking logic. The architecture optimizes the parallel operation efficiency and the resource utilization rate, adapts to the deep learning training and reasoning full scene, and has the advantages of real-time performance and low power consumption.
Owner:GUANGDONG INST OF INTELLIGENT SCI & TECH

High-precision random calculation method based on binary partial product and multiplier

The invention belongs to the field of heterogeneous approximate calculation, and relates to a high-precision random calculation method based on a binary partial product and a multiplier. The high-precision random calculation method based on the binary partial product directly uses the weight of a bit corresponding to a binary number to generate a random bit stream. Bit streams are combined in a uniform distribution mode, phase and sum are carried out bit by bit, and all phase and results are added to obtain the random calculation multiplier with the optimal precision. According to the method, binary weights are combined, a shifting and splicing mode is used for replacing addition, partial product addition is used for achieving random calculation, and an approximate parallel counter is adopted for optimization. The multiplier provided by the invention fully combines the advantages of random calculation and binary calculation, the hardware cost is less than half of that of a binary multiplier, and the delay is lower. When the provided multiplier is applied to neural network reasoning, the precision of model loss can be almost ignored.
Owner:BEIJING INFORMATION SCI & TECH UNIV

Implementation method of dynamic precision approximate multiplication based on partial product decoupling and multiplier

The invention provides a partial product decoupling-based dynamic precision approximate multiplication implementation method and a multiplier, and the multiplier comprises a significance evaluation unit which splits a first operand and a second operand into high digits and low digits, and correspondingly generates a first amplitude and a second amplitude; different precision instructions are generated based on the amplitudes and the threshold values; the dynamic precision unit performs zero setting processing on the second low digit based on the precision instruction to form an approximate number, and calculates a partial product of the first high digit and the approximate number; the accurate calculation unit accurately calculates products of the other three parts; and the final addition unit sums the products of the four parts according to the weight. The calculation unit can be used for constructing or optimizing a multiplication calculation array in a neural network hardware accelerator, is particularly suitable for scenes with dense fixed-point number operation such as convolution, and performs efficient approximate operation on the input fixed-point number through a data driving mechanism by utilizing the common sparsity and amplitude difference characteristics of neural network data, so that the calculation efficiency is improved. And self-adaptive management of power consumption is realized.
Owner:EHIWAY MICROELECTRONIC SCI & TECH (SUZHOU) CO LTD

Multiplier circuit

PendingUS20260186743A1Crossbar switchBinary multiplier
The present description concerns a circuit configured to perform a multiplication between first and second values each comprising up to a number N of chunks of n bits, comprising: a number NMUL of multiplier sub-circuits configured to generate partial products corresponding to a multiplication between one chunk of the first value and one chunk of the second value; an adder part configured to generate an intermediate sum, corresponding to the addition of partial products of same significance, and excess bit values; an input crossbar circuit configured to select at least one chunk of the first value and at least one chunk of the second value and to provide them to one of the multiplier sub-circuits; a control unit configured to indicate which chunks to supply to which multiplier sub-circuit; and a parallel adder configured to generate a value corresponding to the product.
Owner:COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES

In-memory multiply-accumulate and similarity calculation circuit based on dynamic logic

The invention discloses an in-memory multiply-accumulate and similarity calculation circuit based on dynamic logic, which is used for supporting the design of an in-memory calculation macro cell and an accelerator in a digital domain, and provides an in-memory calculation circuit architecture which supports a Booth coding rule for inputting similarity calculation. On the basis, a local input sharing calculation unit and a decoding and addition tree unit are designed, and element-by-element multiply-accumulate calculation and similarity calculation of fixed data are supported. Compared with a traditional multiply-accumulate circuit, the multiply-accumulate circuit disclosed by the invention adopts a calculation mode of firstly carrying out Booth coding on input data, then generating a multiply-accumulate partial product and finally carrying out decoding accumulation, and meanwhile, a partial product generation and similarity calculation circuit is designed based on dynamic logic, so that the partial product is multiplexed, and multiply-accumulate calculation is realized element by element with smaller calculation overhead; the face effect and the energy efficiency are both considered, and the comprehensive performance of the computing architecture and the circuit is improved.
Owner:SOUTHEAST UNIV

An approximate 4-bit lookup table multiplier

This invention provides an approximate 4-bit lookup table multiplier, relating to the field of integrated circuits, comprising: a conversion module, a first-stage multiplication module, and a second-stage multiplication module. During the multiplication process, for two compressed partial products with the same input, a 6-input, 2-output lookup table (LUT6_2) is used for simultaneous calculation. For three compressed partial products with the same input, the precision loss of each compressed partial product is calculated, and a 6-input, 2-output lookup table (LUT6_2) is used to simultaneously calculate the two compressed partial products with the smallest precision loss. The compressed partial product with the largest precision loss is calculated separately using a 6-input lookup table (LUT6). For compressed partial products with an input number exceeding the lookup table's input limit, and for unprocessed partial products, a partial carry-truncation method is used to calculate the product result. This invention can improve multiplier performance while reducing multiplier precision loss.
Owner:YUNNAN UNIV

Zero-frequency signal amplitude reduction method and system based on multi-truncation full-parallel FIR (Finite Impulse Response) filter

The invention discloses a zero-frequency signal amplitude reduction method and system based on a multi-cut full-parallel FIR filter, and belongs to the technical field of chip digital design, and the method comprises the steps: executing the CSD multiplication of a filter coefficient and input data, and obtaining a plurality of partial products; carrying out cutting processing on the plurality of partial products respectively; calculating a compensation value of a truncation error based on the symbol attributes of the plurality of partial products and the state of the least significant bit after truncation; and accumulating the plurality of partial products after cutting processing and the compensation value to obtain filtering output. The invention solves the problems that the zero-frequency signal amplitude of the filter is obviously increased and the resource efficiency and the signal precision are difficult to consider at the same time due to the fact that the truncation digits are increased for reducing hardware resources in the prior art.
Owner:YUNCHIP MICROELECTRONICS CO LTD

Approximate floating point fused dot product step to power circuit

PendingCN122311087AHemt circuitsFloating point
This invention provides an approximate floating-point fused point integral step-by-step alignment circuit, relating to the field of approximate circuit design. The invention generates two sets of effective dot product exponents based on the comparison results, and remaps the inputs of the first and second multiplication modules. The difference between the two sets of effective dot product exponents is applied step-by-step to the arithmetic shift of the approximate partial product, reducing the area of ​​the shift unit and the logic depth. A fine-grained step-by-step shift compression architecture is adopted, integrating alignment and compression operations within two processing units, avoiding large-width shifters becoming critical path bottlenecks. The processed first set of approximate partial products, the second set of approximate partial products after two compressions and shifts of the remaining partial products, and the sign compensation bit are fused and compressed, improving the efficiency of the fusion compression module and effectively reducing circuit area and delay.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Compute-in-memory systems and methods with weight update circuits

A device includes a memory array to store a plurality of weight sets and read circuits to read the plurality of weight sets. A first weight buffer is to store a first weight set of the plurality of weight sets and a write driver circuit is to write the first weight set into the first weight buffer during a single write clock cycle. A plurality of first multiplier circuits receives the first weight set from the first weight buffer and a first data input set of data input channels 0-N. Each of the first multiplier circuits is to receive a corresponding first weight of the first weight set and a first data input of the first data input set and to multiply the first weight and the first data input to provide a partial product. An adder tree is to sum the partial products and provide an accumulated result.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

A compressor-based FPGA approximate multiplier

The application provides a compressor-based FPGA approximate multiplier, and belongs to the field of integrated circuits.The compressor-based FPGA approximate multiplier comprises a partial product generation module, which is used for performing AND operation on each bit of a multiplier and each bit of a multiplicand to generate a partial product matrix; a first accurate compression module, which is used for performing first accurate compression on the partial product matrix based on a plurality of grouping carry compressors to obtain a first compressed element matrix; wherein each grouping carry compressor comprises two LUTs; a second approximate compression module, which is used for performing second approximate compression on the first compressed element matrix by using a plurality of LUTs to obtain a second compressed element matrix; and a carry adder module, which is used for determining a product result by using a carry chain of an FPGA according to the second compressed element matrix.The application improves the utilization rate of an FPGA look-up table and reduces the power consumption of the multiplier.
Owner:YUNNAN UNIV

Method and apparatus for implied bit handling in floating point multiplication

Devices and methods are provided for performing, by a processor in response to a floating point multiply instruction, multiplication of floating point numbers. An example processor includes first, second, third, and fourth computational paths. In operation, the first determines values of implied bits of mantissas of floating point numbers and generates first partial product terms, the second multiplies remainders of the mantissas to generate second partial product terms, the third detects a number of leading zeros in the mantissas and determines a shift amount for each of the mantissas, and the fourth calculates exponents for a flush-to-zero mode.
Owner:TEXAS INSTRUMENTS INC

A cascaded full adder based compressor and partial product array compression method for multiplication operations

The application discloses a compressor based on a cascaded full adder for multiplication operation and a partial product array compression method, and relates to the technical field of computer hardware design. The compressor comprises a first parallel 3-2 compressor circuit, a first parallel cascaded 3-2 compressor circuit, a second parallel 3-2 compressor circuit, a second parallel cascaded 3-2 compressor circuit and a serial 3-2 / 2-2 compressor circuit. The input partial product array is compressed by the first parallel 3-2 compressor circuit, the first parallel cascaded 3-2 compressor circuit, the second parallel 3-2 compressor circuit and the second parallel cascaded 3-2 compressor circuit, and a fourth initial compression result is obtained. Then, the serial 3-2 / 2-2 compressor circuit is used for accumulation, and a target compression result is generated.
Owner:GUANGDONG UNIV OF TECH

Switch for a merchandise forwarder

ActiveCN309802203SMechanical engineeringPartial product
1. The name of the design product: switch of the commodity forwarder. 2. The use of the design product: the use of the overall product is the commodity forwarder which can pass through the elastic pushing force of the spring to make the displayed commodities continuously push forward, and the use of the partial product is the switch of the commodity forwarder. 3. The design points of the design product: the shape of the partial product. 4. The picture or photo which can best show the design points: the perspective view. 5. Other circumstances which need to be explained: the part drawn by the dotted line is the partial product which is not claimed in the present case.
Owner:KAWAJUN KK

High speed low power approximate multiply accumulate operator for image convolution processing

The application discloses a high-speed low-power approximate multiply-accumulate operator for image convolution processing, which processes 8bit signed number*8bit unsigned number. The multiply-accumulate operator is composed of a multiplier and an adder, and is divided into four stages of partial product generation, partial product compression, carry addition and accumulation. In the partial product generation stage, an approximate base-8 Booth algorithm is used, and in the carry addition and accumulation stages, an approximate 4-carry adder is used. The multiplier of ±3 times is approximated to ±2, and the multiplier of ±4 times. Due to the delay of the shift operation, the power consumption is small. Through the approximation, the complexity of the circuit can be reduced, the speed can be improved, and the power consumption can be reduced. For nbit addition operation, n-1bit carry is required; for 16bit addition operation, 4bit carry is used, the delay is reduced, and the calculation speed is improved. Through approximate calculation, the speed of the multiply-accumulate operator is improved, the power consumption is reduced, and the image convolution processing with high speed, low power consumption and certain error tolerance is suitable.
Owner:BEIJING UNIV OF TECH

A hardware accelerator for key switching algorithm in CKKS homomorphic encryption algorithm

The application discloses a hardware accelerator for a key switching algorithm in a CKKS homomorphic encryption algorithm, and belongs to the field of privacy calculation and fully homomorphic encryption hardware acceleration. The hardware accelerator comprises an NTT / INTT circuit and a peripheral polynomial operation circuit. The NTT / INTT circuit is used for realizing domain transformation of a polynomial, and comprises a plurality of parallel butterfly operation cores and an NTT control module for generating a memory address and a data flow control. The butterfly operation core comprises a storage array, a calculation array and a peripheral digital circuit. The storage array is used for storing a rotation factor, and the calculation array is used for calculating a partial product of the rotation factor and polynomial coefficient multiplication. The hardware accelerator can realize highly parallel key switching operation, solves the problem of excessive fully pipelined data flow storage overhead caused by data dependency in other works, and effectively reduces the area and power consumption of the circuit.
Owner:ZHEJIANG UNIV

Low-power variable-width storage-computing integrated array structure

PendingCN122263991AComputations using electromechanical counter-type accumulatorsPhysical realisationParallel computingPartial product
The application discloses a low-power-consumption variable-bit-width memory-computing integrated array structure for one-dimensional convolution operation, wherein when the memory-computing integrated array performs multiplication and accumulation operation, the multiplicand has been written into the array, the multiplier is input in a bit-serial manner, the memory-computing unit completes single-bit multiplication and obtains a partial product; an addition tree in the array completes accumulation of the partial product and obtains a partial sum of a multi-bit unit; a configurable bit-width shift and accumulation module groups, shifts and accumulates partial sums of a plurality of multi-bit calculation units to form partial sums of subarrays; a channel-level addition tree accumulates the partial sums between each subarray to obtain partial sums of the channel-level addition tree; and the partial sums pass through a quantization shift and accumulation module to shift and accumulate the calculation results in a bit-serial manner cycle by cycle to obtain a final calculation result. The application can flexibly adapt to different types of one-dimensional convolution in a light-weighted epilepsy monitoring network, and realizes collaborative optimization of space utilization and computing energy efficiency while ensuring functional correctness.
Owner:PEKING UNIV SHENZHEN GRADUATE SCHOOL

Method and apparatus for multiplication of large-scale matrices by powers of integers modulo 2.

This application provides a method and apparatus for multiplying large-scale matrices by powers of integers modulo 2. The method includes: performing multiple vector inner product operations on each divided block matrix pair, and calculating the modulo multiplication result. The vector inner product operation includes: expanding the partial product based on multiple vector element pairs participating in the vector inner product operation, and retaining the effective partial products with bit weights lower than the target bit width after expansion, according to the target bit width; obtaining the modulo result of a single vector inner product based on the effective partial products; aggregating the modulo results of all vector inner products of the same block matrix pair to obtain the modulo multiplication result of the block matrix pair; and accumulating the modulo multiplication results of all block matrix pairs to obtain the final modulo multiplication result. The method pads high-order bits for partial products whose width is insufficient to reach the target bit width, and truncates partial products whose width exceeds the target bit width, so that subsequent operations only process effective low-order data, thereby reducing hardware resource consumption and the number of addition stages.
Owner:NANJING UNIV

Accelerator optimized using activation function sparsity and computational method of the accelerator

An accelerator, which is optimized by using activation function sparsity, includes a memory and a processor, a Bit-Separable Multiplier (BSM) configured to divide a first data into upper bits and lower bits when first and second data are input, calculate a Partial Product (PP) value for the upper bits by multiplying the upper bits of the first data by the second data, and output a MAC value for the upper bits by performing a multiply-accumulate (MAC) operation, and a dynamic range decoder configured to predict an activation function output value based on the MAC value for the upper bits, and transmit a masking signal to the BSM for instructing to skip the multiplication operation of the lower bits of the first data and the second data when the activation function output value is 0, and set and output an output feature map value to 0.
Owner:KYUNGPOOK NAT UNIV IND ACADEMIC COOP FOUND

Convolution Circuit, Convolution Computing Method, Chip, and Electronic Device

A convolution circuit includes a plurality of multipliers, a first adder coupled to the plurality of multipliers, and a second adder coupled to the first adder. Each multiplier includes a plurality of precoders, a plurality of encoder groups, and an adder tree circuit. Each precoder is in a one-to-one correspondence with one encoder group. Output ends of the plurality of encoder groups and input lines of the adder tree circuit are of a same quantity and in a one-to-one correspondence. In addition, the adder tree circuit is coupled to the first adder. The second adder is further coupled to a memory. A partial product that is related only to a weight parameter may be first accumulated with a constant 1 in the multiplier, and then added to results output by adder tree circuits in the second adder.
Owner:HUAWEI TECH CO LTD