Compact Arithmetic Processing Elements for Low-Precision Parallel Computing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computing architectures inefficiently utilize transistors, providing only a fraction of the computing power available due to a focus on high precision arithmetic, which is not necessary for many applications, leading to wastage of resources.
Innovation Solution
Implementing low precision high dynamic range (LPHDR) processing elements that perform arithmetic operations, using logarithmic or analog representations to maximize the number of operations per unit of silicon area, enabling massively parallel computation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional high precision arithmetic processing elements are used, then measurement precision is improved, but productivity deteriorates
Solution Approach 1:
The patent changes the precision parameter from conventional high precision (32-bit or 64-bit floating point) to low precision (8-bit or 16-bit fixed point), enabling significantly more processing elements to be packed into the same silicon area. This parameter change allows the system to achieve higher computational throughput by executing more operations in parallel, even though each individual operation has reduced precision. The low precision is sufficient for many applications including image processing, audio processing, and machine learning inference.
2Productivity
If more arithmetic processing elements are implemented, then productivity is improved, but device complexity increases
Solution Approach 1:
The patent segments the computational workload across thousands of simple, identical processing elements rather than using a small number of complex elements. Each processing element is designed to be extremely simple (performing basic fixed-point arithmetic), but the collective array of thousands of these simple elements achieves high computational throughput. This segmentation approach simplifies the design of individual elements while maintaining high overall productivity.
Solution Approach 2:
The patent employs partial action by implementing only the essential arithmetic operations needed for specific application domains (such as multiply-accumulate operations for neural networks or convolution operations for image processing). Rather than implementing a complete set of complex arithmetic operations, the processor focuses on a subset of operations that provides sufficient functionality for the target applications, thereby reducing device complexity while maintaining high productivity for those specific tasks.
3Productivity
If low precision arithmetic is used, then productivity is improved, but measurement precision deteriorates
Solution Approach 1:
The patent changes the precision parameter from conventional high precision to low precision (8-bit or 16-bit fixed point), enabling significantly more processing elements to be packed into the same silicon area. This parameter change allows the system to achieve higher computational throughput by executing more operations in parallel, even though each individual operation has reduced precision. The low precision is sufficient for many applications including image processing, audio processing, and machine learning inference.
Data Source
AI summary
A processor or other device, such as a programmable and/or massively parallel processor or other device, includes processing elements designed to perform arithmetic operations (possibly but not necessarily including, for example, one or more of addition, multiplication, subtraction, and division) on numerical values of low precision but high dynamic range (“LPHDR arithmetic”). Such a processor or other device may, for example, be implemented on a single chip. Whether or not implemented on a single chip, the number of LPHDR arithmetic elements in the processor or other device in certain embodiments of the present invention significantly exceeds (e.g., by at least 20 more than three times) the number of arithmetic elements, if any, in the processor or other device which are designed to perform high dynamic range arithmetic of traditional precision (such as 32 bit or 64 bit floating point arithmetic).


