A reconfigurable SIMD vector processing circuit dynamically adjusts bit-width and parallelism through shared multipliers.
Bias generator adds sign-dependent value to intermediate product for fast overflow detection.
Digital signal processor instruction set processes adjacent vector elements using shared arithmetic circuits for real-complex multiplication.
Shared sub-libraries merge common computational functions to reduce logical element consumption, enabling more kernel replication on integrated circuits.
A clock frequency divider uses a shift register to generate non-overlapping phases with precise duty ratios.
A signed multiplier circuit uses a uniform array of programmable logic blocks to build large multipliers, improving design efficiency for programmable ICs.
Segmenting weights into distinct multiplication units sharing a single accumulator reduces power consumption while maintaining calculation accuracy.
A correlithm object processing system aligns linear string objects to perform subtraction operations on real-world numerical values.
Ratio-based conductance normalization in RRAM crossbars performs addition and subtraction, mitigating device non-idealities to maintain computational accuracy.
Staged operators and reducers in a binary look-ahead adder reduce critical path length by avoiding tri-connected structures.
Conversion circuitry performs negative-output conversion of signed digits to unsigned values, eliminating addition circuitry and reducing latency.
A circuit calculates approximate products using truncated mantissa bits to reduce hardware requirements.
Clamping current signals in the linear region reduces variation and improves matching accuracy.
Segmenting multiplication and reading phases via a connectivity mesh reduces discrete operations and energy consumption.
Segmenting the memory array into tiles with bit line selection switches reduces data storage demand while maintaining high computation capability.
A method calculates negative inverse modulus values by segmenting the modulus into groups and processing bits sequentially from least to most significant.
Adjustment circuitry modifies least-significant digits to generate candidate values for efficient rounding.
Reusing a single multiplier in the butterfly operation unit reduces hardware cost and circuit area while maintaining processing efficiency for audio devices.
Radix conversion enables efficient 256-bit multiplication using existing floating-point hardware units.
A data processing device skips zero-weight convolution calculations and selects single values for identical feature maps to reduce power consumption.
Parallel multiplier circuits in a processing pipeline evaluate mathematical functions quickly without increasing device complexity.
Execution unit circuit merges three ALU operations into one cycle via parallel lanes, improving computation efficiency by 33%.
Segmented quantization circuits process specific convolution outputs to reduce errors from wide data distributions, improving neural network precision.
A hybrid hardware multiply-accumulate operator processes floating-point and integer operands to optimize deep learning calculations.
A bit-serial computation system adjusts clock frequency to maintain processing speed under noise conditions.
Coordinate rotation in a segmented architecture reduces power consumption while maintaining calculation precision.
Partitioning the first operand into even and odd portions allows a microprocessor to execute carryless multiplication without generating erroneous carry bits.
A dual-path adder circuit selects between low-power and high-power modes based on input sign and exponent attributes.
A non-volatile memory latch structure performs binary neural network inference using XNOR operations within the page buffer.
Distributed arithmetic circuit uses look-up tables to perform binomial product-sum operations on paired data coefficients.
Decomposing block floating point vectors into reduced bit-width sub-vectors maintains neural network accuracy despite lowered precision processing speed.
A multiplier system partitions operands into subcomponents to compute partial products using smaller processing blocks.
A dynamic data quantization scheme reduces power consumption in convolutional neural network hardware by using shared exponents and reduced bit widths.
Segmenting memory cells into identical words allows valid units to compensate for defects, improving reliability without increasing memory size.
Parallel carry look-ahead adders in the MD5 arithmetic unit resolve processing speed bottlenecks by reducing clock cycles from 256 to 128.
Hybrid Comparison Look Ahead Merge network reduces resource demands while achieving higher streaming throughput via radix pre-sort parallelization.
Interleaving booth encoders with compressor levels in a folded layout reduces wiring distances, lowering propagation delay and power consumption.
A fused multiply-adder uses a Booth encoder and fraction multiplier to process operands across precision levels.