Segmenting parameters into sub-values with varying precision levels reduces energy consumption while maintaining performance.
A programmable look-up table stores linear segment parameters to approximate non-linear activation functions in neural networks.
Optical parametric oscillation in nonlinear materials implements ReLU functions, reducing energy-time product to 1.2×10−27 J s.
A hybrid training method updates optical neural network weights using real signal propagation.
Activity-difference training updates neural network weights using local neuron signals, reducing energy consumption and improving biological plausibility.
Shared operation circuits reduce power consumption while maintaining computational throughput.
Sorting key columns reduces energy consumption by minimizing state transitions in microring resonator accelerators.
A data conversion apparatus transforms neural network structural data to enable high-speed matrix matrix product execution on standard hardware.
Quantizing weights to fixed-point values reduces hardware complexity and energy consumption while maintaining calculation precision.
A processing unit divides and compresses neural network coefficient matrices using flexible range settings in non-filter dimensions.
A neural network unit employs a multiplexed register rotater to shift data rows across processing elements.
Separating real and dummy node parameters into different storage areas prevents unauthorized inference when hackers access partial model data.
Applying sparsity parameters to neural network coefficients reduces memory footprint and computational demands while maintaining measurement precision.
Automated post-training quantization model selection using indirect metrics to identify optimal configurations.
A scheduling framework segments deep learning operations across multiple GPU cores to accelerate computational throughput.
Neural processing units recompute skip connection tensors from input data and layer weights to reduce memory footprint.
A neural network model generates lookup tables in volatile memory during circuit simulation reads.
Semi-federated data selection reduces upload traffic and training time by filtering edge server inputs before retraining deep learning models.
Runtime software issues a frequency control signal to synchronize multiple accelerators, reducing idle times and optimizing power consumption.