Multiple identical AI dies are linked in one chip package so neural network layers can be mapped flexibly while cutting ASIC design time and NRE costs.
Separating Fe-RAM memory and matrix compute dies in a 3D AI chip cuts data-transfer latency while raising bandwidth and lowering energy use.
Separating weight storage and compute across memory and compute dies cuts AI matrix multiplication latency and power with high-bandwidth FeRAM access.
Using identical AI dies in one chip package maps neural network layers efficiently while reducing ASIC design time and non-recurring costs.
Multiple identical AI dies are linked in one chip package so neural network layers can be split across reusable hardware, cutting ASIC design time and cost.
Separating memory and compute dies with PE-based matrix multiplication cuts AI latency and power while boosting bandwidth for training and inference.
Jacobian-based perceptual hardware measures compressed image quality more like human vision while enabling real-time, lower-power processing.
Jacobian mapping into a non-Euclidean perceptual space cuts compression entropy while aligning image and video quality with human perception.
DNB coding removes the ANBD input constant to cut runtime overhead and hardware demand while preserving error detection in coded data processing.
Expected-versus-actual matrix sums detect and correct computation errors, enabling faster, lower-power processors with reliable results.
Blockwise compression removes leading bits and truncates least significant bits to cut processing time and power while preserving needed precision.
Programmable duty cycles in charge-transfer summing circuits cut area and reduce interference when adding static or dynamic multi-channel signals.
A split-path shifter uses inverted exponent bits to avoid bias subtraction and mantissa negation, shortening the critical path.
Multiplication and division in invertible flow layers improve probability estimation and raise lossless compression efficiency.
A reconfigurable FPGA DSP block uses shared multipliers and selectable precision to handle ML and signal processing with high density and lower power.
Direct exponent-bit shift control speeds floating-point to fixed-point conversion while avoiding bias subtraction and mantissa negation.
Current leakage switching routes transistor load paths between circuit nodes to process ternary values and speed neural network inference.
Hardware converts posit bit strings to analog for neuromorphic operations, then back again to improve precision, speed, and power use.
Non-uniform switch timing matched to capacitor settling cuts vector-matrix multiplication cycles, area, and power while preserving accuracy.
Multiplexers and bit shifting let one division/modulo circuit handle multiple power-of-two inputs, cutting chip area and latency.
Arithmetic coding compresses smart contract circuits into compact bit streams, cutting storage and transmission load without losing exact reconstruction.
Scaling neural network weight matrices by thresholds keeps CIM output within ADC range and improves quantization precision.
A domino-logic output region combines LUT outputs to cut device count while preserving speed and lowering area and power.
Converting real numbers to LNS-coded integers before arithmetic coding cuts integer-operation overhead in functionally safe processing.
Periodic modulo behavior compresses RNS lookup tables into linear arrays, cutting memory for parallel addition, subtraction, and multiplication.
Flexible-precision FPGA DSP blocks use efficient weight loading and parallel multiply-accumulate paths to handle AI and DSP tasks with lower power.
Processes AV1 pre-carry words on the fly using hold-sum and FFcount logic to cut buffer memory, latency, and hardware cost.
Direct radix-4 decoding and tripler circuitry cut multiplier area and energy use while improving partial product generation for AI chips.
A configurable DSP block combines tensor processing, bandwidth expansion, and selectable arithmetic modes to handle ML and DSP tasks efficiently.
Concurrent ECC and multiplication in a PIM MAC pipeline correct stored-data errors on the fly, preserving AI compute accuracy without added delay.
Zero-bit-heavy CNN weights are kneaded into compact segments, cutting invalid convolution work, latency, and power use on lightweight devices.
A shared division and modulo circuit uses multiplexers and bit shifting to process power-of-two input sets with less chip area and latency.
Inductor-assisted folded cascode amplification preserves PAM4 linearity at high frequency without raising supply voltage or shutting down transistors.
Segmented weight-bit processing computes binary scalar products with simpler ADCs, removing DAC overhead and reducing analog error and energy use.
Grouped encoders and inversion compressors cut MOS transistor count in a multiplier, reducing area and power while keeping compression accuracy.
Smaller scan transistors create a delay choke in an SDFQ multiplexer, holding scan signals long enough to avoid race conditions.
A reconfigurable DSP block changes weight-loading modes to support both machine learning and signal processing with high density and lower power.
Single-instruction hardware conversion scales and reformats values directly, cutting intermediate steps, cycles, and resource use.
Parallel ECC and MAC execution corrects data errors during in-memory multiplication, improving AI calculation speed and accuracy.
Sparse reference arrays match CIM temperature coefficients at the ADC, reducing drift without op-amp power or speed penalties.
Non-uniform switch timing matched to capacitor settling cuts vector-matrix multiplication cycles, space, and energy while preserving accuracy.
A reconfigurable DSP block changes weight loading paths to support both AI tensor math and DSP operations with high density and lower power.
A revised recursive network shortens critical paths for higher sinusoid output rates while periodic refresh limits quantization error.
Analog conversion lets posit-stored bit strings run neuromorphic operations faster while preserving precision and lowering power use.
Lossless arithmetic coding compresses serialized arithmetic circuits, cutting blockchain storage and bandwidth while preserving smart contract execution.
A shifter-based circuit converts floating-point values to fixed-point format without exponent bias subtraction, cutting conversion time and hardware.
A configurable FPGA DSP block uses flexible precision and efficient weight loading to handle ML and signal processing on one chip.
Shared weight registers and parallel multipliers let one FPGA DSP block handle AI and digital signal processing with better density and power use.
Arithmetic coding compresses serialized smart-contract circuits to cut blockchain storage and bandwidth while preserving exact reconstruction.
Min-max lossy tensor compression increases effective on-chip memory capacity and bandwidth for neural network inference while preserving accuracy.
Redistributing partial product bits across PLD carry chains and ternary compression cuts routing, area, and power in soft multipliers.
Switched capacitors and SRAM perform digital-input multiplication by charge sharing, cutting data transfer, clock cycles, and energy use.
Dual accumulators downsample high-rate ADC samples so software can run slower while preserving resolution and slope estimation.
Direct LUT-to-adder links and a sparse interconnect structure enable ternary and binary addition in PLDs without added routing overhead.
Opcode-controlled multiplexers and adders let one ALU handle concurrent DSP arithmetic and logic without PLD fabric bottlenecks.
A 3D cross-point memory array executes sum-of-products operations using programmable conductances at cell cross-points.
A Pade approximation convert circuit generates sinusoidal wave signals using a multiplier, divider, and adder.