Fusing depthwise and pointwise convolution kernels within a compute-in-memory array to process pre-activations directly in memory.
Half adder logic generates carry bits for most significant positions, enabling rounding increment injection before final summation to reduce cycle time.
A carry-save multiply-and-accumulate unit processes signals in carry-save format to reduce propagation delay.
A computational system performs Newton-Raphson iterations to produce results with additional accuracy bits for subsequent analysis.
A truncated summation array performs division by discarding less significant columns to reduce circuit area.
Segmented latch circuits synchronize odd and even data paths to resolve accumulation speed bottlenecks in deep learning MAC operations.
A neural network device uses convolution SRAM and diagonal accumulation SRAM to process sparse weight matrices efficiently.
A multiply-accumulate unit sums signals after parasitic capacitor transients stabilize.
Task entanglement-based coding encodes distributed matrices using Chebyshev polynomials to enable partial result recovery from edge devices.
An ontology decomposes user requirements into component-level functionalities to automate service configuration generation.
Expansion markers route composite data elements to standard computing operations, eliminating the need for specialized processing blocks.
Associative memory arrays perform concurrent bit comparisons to replace sequential subtraction operations.
A data processing method splits large numbers into segments for efficient multiplication using register operations.
Dimension transformation matrices reduce model degrees of freedom, accelerating thermal fluid simulation while maintaining solution accuracy.
External zero detectors skip multiplications for zero weights in systolic arrays, reducing dynamic power consumption while increasing device complexity.
A parallel calculation method computes look-ahead parameters using approximated intermediate results to accelerate modular multiplication operations.
Integrating 2's complement completion with rounding operations reduces 1's complement delay, improving FMA operation speed.
Direct Boolean formula interpretation via truth table analysis eliminates CNF conversion overhead, enhancing hardware solver performance and capacity.
A divider module calculates division by multiplying the dividend with a reciprocal binary fraction derived from the divisor.
A systolic array processes complex matrices directly using modified Givens rotations for MIMO decoding.
Segmented accumulators reduce power consumption and precision loss in neural network computations.
A multiply-accumulate circuit uses time-domain delay chains to perform neural network calculations with reduced hardware complexity.
Shuffled secure multiparty deep learning segments randomized data parts across samples to enable external computation without reconstructing original information.
A multioperand decimal adder uses binary carry-save addition to process multiple BCD operands directly.
A processing system for binary weight convolutional neural networks replaces multiplication with addition and subtraction operations.
Eliminates dedicated carry chains by configuring FPGA lookup tables in normal mode, enabling post-mapping optimization and reducing area usage.
An internal data handler stores operands in registers, reducing memory access and power consumption while minimizing exposure to external attacks.
A method distributes error bounds across operators in a data-flow graph to optimize hardware logic implementation.
Time-delay based arithmetic replaces floating-point units to resolve the trade-off between calculation speed and power consumption in deep learning hardware.
Partitioning datasets across 3D-stacked memory banks reduces data movement overhead and random accesses during in-memory radix sorting.
A processor converts divisor reciprocals to fixed point for GPU division.
Lookup table memory stores seeds for constant multiples to generate partial products, reducing computational complexity in machine learning.
Sensing circuitry executes arithmetic operations on bit strings within the memory array.
Capacitive charge accumulation replaces continuous current flow to reduce power dissipation and increase resolution beyond RRAM limitations.
A carry propagate adder stage processes non-final intermediate operands to generate final sums or limit values.
A memory-centric neural network hardware accelerator updates weight matrices using timestamp registers and lookup tables for efficient processing.
A digital image quantization method replaces division with scaled reciprocal multiplication operations.
Segmenting 32-bit operations into 16-bit steps reduces power consumption and silicon area in low-power CPUs.
Group algebra representations map matrices to vectors for recursive block-diagonal multiplication, resolving O(n^3) computational bottlenecks.
A predictive adder circuit generates results for incremented and decremented sums by identifying ripple bit patterns in parallel with the primary addition operation.
Shared arithmetic circuits switch weight data via context signals, reducing circuit size and power consumption while maintaining recognition accuracy.
A digital data processor converts N-bit signals to M-bit outputs using weighted addition and arithmetic rightward shift circuits.
Shared alignment and normalization stages in a fused floating-point adder-subtractor reduce operational latency and circuit area requirements.
Dynamic register output modes eliminate dedicated transpose hardware, reducing device complexity while maintaining matrix multiplication efficiency.
Floating point polynomial evaluation partitions input domains into sub-domains to assign location-based precision, bounding relative error below one.
Separate carry entrance and exit paths isolate buffers from inputs, reducing capacitive loading and propagation delay.
A mass multiplier circuit processes discrete values simultaneously to increase computational throughput.
A two-phase algorithm computes Gaussian integer congruents using bit shifts and truncations to replace costly division operations.