Configurable-Precision AI Processing Elements for Low-Power NPU Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing electronic devices face challenges in efficiently performing artificial neural network inference operations due to high power consumption, heat generation, memory requirements, and cost, particularly when implementing neuromorphic integrated circuits, which are prone to noise and require analog-to-digital conversion, making them unsuitable for compact designs.
Innovation Solution
A standalone, low-power, low-cost neural network processing unit (NPU) with a digital processing element array, SRAM memory, and an NPU scheduler that optimizes resource allocation and memory reuse, allowing for efficient inference operations across various devices by minimizing power consumption and memory usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If neuromorphic integrated circuits are used for neural network inference, then inference capability is provided, but power consumption increases and noise susceptibility worsens
Solution Approach 1:
The system segments neural network processing into distinct functional units: digital signal processing for inference operations, separate analog-to-digital conversion stages, and modular neural network model storage. This segmentation allows digital components to handle computation while minimizing analog circuit exposure, thereby reducing power consumption and noise susceptibility.
Solution Approach 2:
The patent replaces analog neuromorphic circuits with digital signal processing systems. Digital logic circuits perform neural network inference operations instead of analog circuits, eliminating the inherent noise and high power consumption associated with analog implementations while maintaining inference capability.
2Measurement precision
If analog-to-digital conversion is implemented in neuromorphic circuits, then signal processing capability is improved, but device complexity and area increase
Solution Approach 1:
The patent extracts and separates the analog-to-digital conversion function into distinct conversion stages, removing it from the core neuromorphic circuitry. This allows the analog conversion to be performed only when necessary, reducing the overall complexity and area of the integrated circuit while maintaining signal processing capability.
Solution Approach 2:
The system introduces digital signal processing as an intermediary between analog input signals and neural network computation. This intermediary layer handles signal conditioning and conversion efficiently, reducing the complexity burden on the main neuromorphic circuit while preserving measurement precision.
3Measurement precision
If memory resources are increased for neural network models, then inference accuracy is maintained, but power consumption and cost increase
Solution Approach 1:
The system performs preliminary processing of neural network models to optimize their representation for efficient storage and retrieval. By preprocessing models into compressed or optimized formats before loading them into memory, the system maintains inference accuracy while reducing the memory resources required, thereby lowering power consumption.
Solution Approach 2:
The patent dynamically adjusts memory allocation parameters based on the specific neural network model being executed. By changing memory usage parameters adaptively rather than allocating fixed high-capacity memory, the system maintains the precision needed for accurate inference while minimizing power consumption and cost.
4Reliability
If digital processing element array is used instead of analog circuits, then noise susceptibility is reduced, but manufacturing complexity increases
Solution Approach 1:
The digital processing element array is designed to perform multiple functions: neural network inference operations, signal processing, and control logic. This multi-functionality reduces the need for separate specialized circuits, thereby simplifying manufacturing while maintaining noise immunity through digital architecture.
Data Source
AI summary
A neural network processing unit (NPU) includes a plurality of processing elements, each processing element comprising at least a multiplier configured to receive weight parameters of a first predetermined bit-width and input activation data of a second predetermined bit-width, an adder, an accumulator for performing multiply-accumulate (MAC) operations, and a bit quantization unit for generating output activation data of a third predetermined bit-width.


