Waveguide reflection and nonlinear electro-optic conversion replace time-delay feedback, enabling faster, more scalable reservoir computing.
A policy-based neural model predicts joint multi-agent trajectories with scene consistency while cutting memory and computing load.
Electric-field-driven domain wall control in a segmented MTJ enables leaky-integrate-reset behavior plus global inhibition without hard magnets.
Grouped policy networks and Gibbs sampling generate scene-consistent multi-agent trajectories with lower compute and better interaction accuracy.
RIS reflections emulate CNN convolution over the air, enabling low-power real-time inference on IoT devices without onboard processors.
Iterative layer conversion, random initialization, and fine-tuning transfer object recognition knowledge to smaller vehicle-ready neural networks.
A SiON diffusion barrier with oxygen and nitrogen gradients suppresses oxygen penetration and stabilizes threshold voltage to widen memory window.
A shared intermediate layer lets a triboelectric generator drive a synaptic transistor without extra circuits or external power, reducing weight and complexity.
A three-layer stacked light sensor separates pixels, logic, and memory to run neural image processing on-chip while limiting noise.
A paraelectric gate-interposed layer lowers gate-stack capacitance in a FeFET, widening the memory window and improving reliability.
A neural network filters potential maneuvers through a safe, legal action mask to improve autonomous driving transparency and training time.
Ferroelectric field modulation enables one photodetector to switch photocurrent from positive to negative, integrating sensing, memory, and fast image processing.
A floating-body transistor detects light and directly generates spike signals, cutting sensor-chain cost, delay, and power in artificial vision.
A tunneling and ferroelectric layer stack suppresses leakage current and power dissipation while preserving selector access in crossbar memory arrays.
A multi-layer neural network predicts individual Mini-LED transfer errors in advance, improving placement accuracy without slowing mass transfer.
An LSTM time-series model predicts hot-rolling roll-bending force more accurately while reducing input data burden and handling nonlinear coupling.
Hardware accelerators on factory-floor edge devices run neural network matrix operations fast enough to update digital twins in real time.
Two stochastically rounded binary neural networks are combined to preserve classification accuracy while cutting memory use and energy.
A superconducting neuron circuit uses flux accumulation and adjustable thresholds to cut area and power while avoiding synchronous logic operation.
Splitting each N-bit output into conditioned first and second halves cuts sequential RNN computation while preserving output quality on mobile devices.
Splitting N-bit output generation into two conditioned halves cuts sequential RNN computation while preserving accuracy on resource-limited devices.
A dual-mode ASIC staggers tile inputs with delay registers to limit power spikes and current changes during neural network computation.
Inverting circuits smooth pulse-like synapse outputs into sigmoid signals, enabling multi-level neuromorphic processing.
Cross-discrimination between local generative and discriminative models helps a central node aggregate heterogeneous federated models without sharing raw data.
Decoupled weight decay and reduced-precision momentum storage let large Vision Transformers improve downstream accuracy with lower memory and compute overhead.
Selective normalization and rounded right shifts preserve fractional precision in quantized neural networks while cutting memory and compute overhead.
Self-generated calibration data and layer sensitivity analysis shrink DNN models for GPU, DSP, and NPU inference without retraining.
Bidirectional links between co-activated BMUs enable unsupervised multimodal classification with distributed processing that avoids labeled data and central bottlenecks.
Train with expanded blocks, then collapse them into regular convolutions to improve NPU utilization without sacrificing accuracy.
Scheduler-driven PE allocation lets one NPU run multiple ANN models in parallel or time division, cutting idle time, delay, and power use.
By selecting only non-zero input features and weights for convolution, this case cuts redundant neural network operations and speeds real-time processing.
Local on-chip memory and parallel neural elements speed AI inference while cutting the power draw of CPU and GPU processing.
Mixed-precision quantization keeps sensitive weight columns at higher precision to cut LLM GPU memory use with minimal inference loss.
Sparse pairing of weights and activations skips irrelevant neuron operations, cutting ANN power use and hardware load while preserving accuracy.
By splitting tensors and summing transformed sub-results, this case reduces convolution multiplications, power overhead, and hardware latency.
Mode-controlled buffer swapping lets a CNN accelerator process 1x1 and 1xN kernels with fewer fetches and higher MAC utilization.
Periodic optical target detection triggers control only when needed, cutting neural network processing load and power use.
A retinomorphic array with memristor crossbar preprocessing cuts redundant visual data transfer, reducing latency, bandwidth load, and power use.
Compression with pruning, quantization, and visual comparison reduces neural network size and computation while preserving accuracy.
Maps argmax and argmin to existing NNA operations, avoiding off-chip processing and extra hardware while preserving speed and resource use.
A deterministic tie-breaker layer resolves near-tied class probabilities to keep deep learning inference consistent across hardware-software stacks.
Splitting input tensors into distributed tiles enables parallel AI computation while easing memory bottlenecks and limiting data sharing.
Graph traversal identifies redundant neural network channels, cutting coefficients, compute load, and memory bandwidth without changing output.
Tailored model parameters let communication nodes use spare computing capacity for AI learning while adapting to local information and resource limits.
Simultaneous matrix multiplication and All-Gather across partitioned LLM layers cuts multi-device inference latency and communication overhead.
Hierarchical fine-grained sparsity constrains zero patterns so neural networks run faster and use less energy without custom hardware for each sparsity level.
Block-diagonal weight remapping changes group definitions so AI chips run grouped convolution faster with less parallel-processing overhead.
Split neural network layers between sensor devices and an aggregator to cut bandwidth and latency while preserving real-time determination accuracy.
A single-chip logic-gate neural network replaces sequential architectures to deliver real-time inference with minimal power use.
Balances local layer processing and cloud offloading by comparing compute and transmission energy in distributed neural networks.
An Al2O3 insertion layer stabilizes ECM and VCM filaments in a polymer memristor, enabling gradual conductivity updates and better state retention.
Weight preloading, delayed column transfer, and broadcast inputs raise PE array utilization in depth-wise convolution while cutting memory reads and power.
Compensation matrices correct quantization error in neural networks, preserving inference accuracy while keeping memory use within limits.
Threshold values are trained from circuit-model environment differences to reduce quantization errors in neural network inference.
Controlled neural noise synthesis obscures transmitted gradients to block image reconstruction attacks while preserving distributed training fidelity.
Per-layer CNN mode selection cuts external memory bandwidth by switching among input-fixed, output-fixed, and kernel-fixed sliding schemes.
Specialized embedding, attention, and base dies cut AI power and latency by reducing data movement through high-bandwidth EMIB links.
Adaptive residual and attention units on a local neural network processor cut resource use while speeding large-vector AI data processing.
Compressed attention layers use tensor sub-matrices to reduce memory bandwidth, latency, and power demands in hardware neural networks.
Polynomial expansions and Hadamard products replace quadratic attention to reduce compute, memory use, and latency in ML models.
Pre-training measures quantization sensitivity to assign mixed bit widths, reducing memory and compute demands while preserving model accuracy.
A configurable activation module uses localized lookup data to limit silicon area and power while covering changing input ranges.
Separate memory bank arrays store decoder weight matrices and key-value pairs to reduce access collisions and speed output-token generation.