Mixed-Precision Machine Learning Circuits for Layer-Specific Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Integrated circuits face challenges in efficiently performing machine learning tasks due to the need for high-precision inference in certain layers, which can be expensive and inefficient in terms of area or throughput, while operating at lower precision may not be sufficient for all layers.
Innovation Solution
The integrated circuit device employs dynamically mixed precision by using block floating point formats to operate at lower precision for some layers and higher precision for critical layers, utilizing conversion circuitry and splitter components to adjust precision dynamically based on the layer requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all layers execute at high precision, then measurement precision is improved, but device complexity and resource usage increase
Solution Approach 1:
The patent applies different precision levels to different layers of the machine learning graph based on their specific requirements. Critical layers that require high precision for accuracy use higher precision arithmetic, while less critical layers use lower precision to reduce computational overhead. This local differentiation resolves the contradiction by providing high precision only where needed rather than uniformly across all layers.
Solution Approach 2:
The system dynamically adjusts precision levels during execution based on layer requirements and computational needs. The precision mode can be changed between layers, allowing the system to adapt precision to match the specific computational demands of each layer, thereby optimizing the balance between accuracy and resource consumption.
2Device complexity
If all layers execute at low precision, then device complexity is reduced, but measurement precision deteriorates for critical layers
Solution Approach 1:
The system identifies critical layers that require high precision and applies higher precision arithmetic only to those specific layers rather than uniformly across all layers. This localized approach ensures that measurement precision is maintained where it matters most while keeping overall computational resource usage lower.
Solution Approach 2:
Instead of applying high precision to all layers (excessive action), the system applies high precision only to the extent necessary for critical layers (partial action). This selective application of precision resources optimizes the balance between maintaining required accuracy and reducing overall computational complexity.
3Measurement precision
If high precision is used for critical layers, then measurement precision is improved, but area and throughput are worsened
Solution Approach 1:
The patent implements different precision characteristics in different layers, allowing critical layers to use high precision for accuracy while non-critical layers use lower precision that enables higher throughput. This spatial differentiation of precision quality resolves the contradiction between precision and productivity.
Solution Approach 2:
The system dynamically configures precision levels based on layer requirements, enabling high precision where needed for accuracy while maintaining lower precision in other layers to preserve throughput. This dynamic adaptation allows the system to optimize both precision and productivity according to specific computational needs.
Data Source
AI summary
Systems, methods, and circuitry for dynamically mixed precision machine learning are provided. An integrated circuit may include conversion circuitry to convert input feature data to block floating point format and upper/lower splitter circuitry to split the input feature data in the block floating point format into an upper component in the block floating point format and a lower component in the block floating point format. A processing element may use only the upper component when operating in a lower-precision mode and use both the upper component and the lower component when operating in a higher-precision mode.


