Asynchronous Analog MAC Accelerator for Low-Power Edge Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital MAC engines for machine learning applications are power-hungry, expensive, and less secure, making them unsuitable for edge computing and sensor applications that require low power consumption, low cost, and high safety and privacy standards.
Innovation Solution
The development of analog current-mode multipliers and multiply-accumulate (MAC) ICs that operate in current mode, optimizing cost-performance by arranging activation and weight signals according to a programmable statistical distribution, avoiding signal overflow, and leveraging asynchronous, approximate computation to reduce power consumption and silicon area, while eliminating the need for intermediate converters and enabling Compute-in-Memory operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If digital MAC engines are used for machine learning applications, then programmability and computation speed are improved, but power consumption and cost increase significantly
Solution Approach 1:
The patent replaces digital computation systems with analog computation systems. The analog MAC engine uses continuous physical quantities (voltages, currents) to perform multiply-accumulate operations directly, eliminating the need for digital logic circuits, clocks, and memory access operations. This substitution of computational paradigm dramatically reduces power consumption while maintaining the ability to perform machine learning computations.
Solution Approach 2:
The patent changes the operational parameters from digital discrete states to analog continuous states. By using analog voltages and currents that can vary continuously, the system performs computations in parallel without sequential clock cycles, reducing power consumption. The analog signals represent data and weights directly, enabling efficient matrix multiplications for neural networks.
2Measurement precision
If digital MAC engines are used for machine learning applications, then computation precision is improved, but cost and power consumption increase
Solution Approach 1:
The patent employs analog circuits that can be implemented using standard CMOS fabrication processes without requiring expensive advanced nodes. The analog MAC engine uses simple transistor-based circuitry that is cost-effective to manufacture at scale, sacrificing some of the extreme precision of cutting-edge digital circuits in exchange for significantly lower manufacturing costs and power consumption.
3Power
If digital MAC engines are deployed in cloud, then computation power is improved, but security and latency are worsened
Solution Approach 1:
The patent segments the machine learning computation functionality from centralized cloud infrastructure and embeds it directly into edge devices and sensors. The analog MAC engine enables local inference capabilities, dividing the computational workload between edge devices and cloud, thereby improving security by keeping sensitive data processing local while maintaining access to cloud-based model updates.
4Use of energy by moving object
If analog current-mode multipliers are used, then power consumption is reduced, but signal overflow issues arise
Solution Approach 1:
The patent implements dynamic range management in the analog MAC engine through adaptive biasing and signal normalization techniques. The system dynamically adjusts operating points and signal levels to prevent overflow conditions while maintaining low power consumption. This dynamic adaptation allows the analog circuits to handle varying input ranges without losing information or requiring excessive power headroom.
Data Source
AI summary
Methods of performing mixed-signal/analog multiply-accumulate (MAC) operations used for matrix multiplication in fully connected artificial neural networks in integrated circuits (IC) are described in this disclosure having traits such as: (1) inherently fast and efficient for approximate computing due to current-mode signal processing where summation is performed by simply coupling wires, (2) free from noisy and power hungry clocks with asynchronous fully-connected operations, (3) saving on silicon area and power consumption for requiring neither any data-converters nor any memory for intermediate activation signals, (4) reduced dynamic power consumption due to Compute-In-Memory operations, (5) avoiding over-flow conditions along key signals paths and lowering power consumption by training MACs in neural networks in such a manner that the population and or combinations of multi-quadrant activation signals and multi-quadrant weight signals follow a programmable statistical distribution profile, (6) programmable current consumption versus degree of precision/approximate computing, (7) suitable for ‘always-on’ operations and capable of ‘self power-off’, (8) inherently simple arrangement for non-linear activation operations such as Rectified Linear Unit, ReLu, and (9) manufacturable on main-stream, low cost, and lagging edge standard digital CMOS process requiring neither any resistors nor any capacitors.


