Quantized Neural Network Circuit Energy Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep neural networks (DNNs) consume significant power and energy, especially in battery-powered devices, necessitating improvements in energy efficiency for applications like speech recognition, image classification, and robotics.
Innovation Solution
A system with binary neurons and multiplexers in a scalable SIMD architecture, allowing variable precision and energy-efficient quantized neural network inference, using a resource-constrained scheduling algorithm to map operations onto processing elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If MAC-based designs are used for neural network inference, then computational accuracy is maintained, but energy consumption is high
Solution Approach 1:
The patent applies quantization to change the precision parameters of neural network computations from high-precision floating-point (MAC operations) to low-precision integer representations. This parameter change reduces the energy consumption of computational operations while maintaining sufficient accuracy for inference tasks, directly resolving the contradiction between energy efficiency and computational accuracy
Solution Approach 2:
The patent replaces the traditional MAC (Multiply-Accumulate) mechanical computation system with a neural network inference system that uses quantized representations and specialized hardware architectures. This substitution eliminates the need for energy-intensive floating-point multiplications and accumulations, achieving up to 50x energy efficiency improvement while preserving inference accuracy
2Measurement precision
If high precision neural network operations are performed, then accuracy is maintained, but power consumption increases
Solution Approach 1:
The patent changes the precision parameters from high-precision floating-point arithmetic to low-precision integer quantization. This parameter transformation reduces the bit-width of data representations and computational operations, thereby reducing power consumption while maintaining inference accuracy within acceptable thresholds
Solution Approach 2:
The patent applies quantization that uses fewer computational resources than traditional MAC operations, performing only the necessary computations at reduced precision. This partial action approach maintains sufficient accuracy for inference while significantly reducing the excessive power consumption associated with full-precision operations
3Use of energy by moving object
If variable precision quantization is implemented, then energy efficiency improves, but device complexity increases
Solution Approach 1:
The patent segments the neural network computation into multiple stages with different quantization precisions. By dividing the computational workflow into segments that can use different precision levels, the system achieves variable precision quantization benefits while managing device complexity through modular organization of computation stages
Solution Approach 2:
The patent designs a universal quantized neural network inference system that can handle multiple precision levels and operation types using the same hardware architecture. This multi-functionality approach allows variable precision quantization without proportionally increasing device complexity, as the same circuit structures serve multiple precision requirements
Data Source
AI summary
A quantized neural network circuit. The circuit may include a neuron processing element, the neuron processing element including a first neuron cluster and a second neuron cluster. The first neuron cluster may include: a first binary neuron, having a first input network with a first number of inputs; a second binary neuron, having a first input network with a second number of inputs, the second number being different from the first number; a plurality of multiplexers, each having an output connected to a respective input of the inputs of the first input network of the first binary neuron; and a plurality of flip-flops, each having an output connected to an input of a respective multiplexer of the plurality of multiplexers.


