Sub-vector multiplexing and merge stages compress valid vector elements while cutting wire congestion and processor area in AI chips.
Programmable adder hardware combines fixed-width dot products into target precisions, cutting matrix multiplication cost, latency, and footprint.
Programmable adder hardware combines low-bitwidth dot products into target precisions, cutting matrix multiplication cost and hardware footprint.
Quantized binary weights and iterative tuning cut ReRAM overlap variation errors, improving in-memory RNN MAC accuracy for edge hardware.
A unified bit mapping and permutation scheme simplifies Polar code hardware across code lengths while improving decoding and lowering block error rate.
Converting FP32 inputs into TF32 operands lets GPUs run matrix multiply-accumulate faster with lower energy use and latency.
Embedding-based bitstream selection predicts co-execution at network nodes, cutting cloud FPGA latency without redundant cache storage.
Splitting vector data across multiplexer sets and merging sub-vectors later reduces wiring density, processor area, and delay.
Dedicated NIC reformat circuitry converts data in transit with a single read and write, cutting CPU load and memory bandwidth waste.
Single-ECC sectors are split into spaced track portions to improve HDD reliability under local impairments without read-modify-write penalties.
A tiled switch matrix uses configurable multi-block switching stages to permute data patterns efficiently without many custom circuits.
Cyclic shift units and cache reordering enable streaming 2D and 3D matrix transforms with lower circuit complexity and power use.
Adjacent switching blocks and staged cross-block routing enable flexible data permutations while separating control and data paths for higher bandwidth.
An iterative seed chain with permutation and bijective mapping generates pseudo-random codes on the fly with lower hardware complexity and power.
Embedding-based bitstream selection improves cloud FPGA cache hits by prefetching configurations likely to run together across tenants.
Cross-block switching stages let one matrix remap input data into multiple patterns without extra control logic for dynamic AI workloads.
Direct arithmetic on order-preserving compressed data avoids decompression bottlenecks, cutting latency, throughput loss, and energy use.
Maps input bits by M_index so one Polar coding hardware path can support multiple mother code lengths with better BLER and 5G efficiency.
Multi-cycle generation of w-bit subsets enables flexible polar block conditioning with lower hardware usage and shorter clock cycles.
Parallel bit pattern generation cuts polar coding clock cycles while supporting flexible K, N, and M settings with lower hardware usage.
Converting bit strings between floating-point and posit formats lets computing tiles raise arithmetic precision without increasing memory demands.
Cyclic shift stages and cache reordering enable flexible matrix transposition and dimension expansion with lower hardware resource use.
Spanning switching stages let a tiled switch matrix remap data patterns dynamically across adjacent blocks without extra control logic.
Combined bit-reversal and memory transpose in an FFT engine cuts latency and circuit area for high-throughput spectral processing.
Transform matrices replace length-specific Polar encoding hardware, simplifying 5G NR sequence determination and improving resource reuse.
Operate on order-preserving compressed data without decompression to cut processing load, energy use, and bandwidth.
Bit-reversed 2-D mapping with pseudo-random column rotation disperses periodic interference and burst noise across OFDM bandwidth.