Hybrid Accumulation Method for MAC in Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning integrated circuits (ICs) for edge-based sensors face challenges with high power consumption, cost, latency, and security concerns due to reliance on cloud-based processing, which is not suitable for real-time applications like medical devices and requires a shift towards low-cost, low-power, and flexible solutions for local computation.
Innovation Solution
The development of MAC ICs with spatial-temporal multiply-accumulate operations, hybrid signal processing, and mixed-signal accumulators that reduce overflow constraints, power supply limitations, and cumulative errors, enabling asynchronous operation and low-power consumption while maintaining programmability and flexibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If cloud-based processing is used for machine learning operations, then computational power and flexibility are improved, but power consumption, latency, and security risks increase
Solution Approach 1:
The patent segments the machine learning processing into two parts: complex training operations remain in the cloud, while inference operations are performed locally on edge devices. This segmentation allows the system to leverage cloud computational power for model development while minimizing ongoing power consumption by performing only lightweight inference locally.
Solution Approach 2:
The patent introduces an intermediary layer (edge device with local MAC operations) between the cloud and end applications. This intermediary performs local processing of commonly accessed data patterns, reducing the need for continuous cloud communication and thereby reducing power consumption while maintaining computational effectiveness.
2Power
If cloud-based processing is used for machine learning operations, then computational capabilities are improved, but latency increases
Solution Approach 1:
The patent performs preliminary actions by pre-processing and caching commonly accessed data patterns locally on edge devices. This preliminary local processing eliminates the need for real-time cloud communication during inference, significantly reducing latency for time-critical applications while maintaining computational accuracy.
3Measurement precision
If cloud-based processing is used for machine learning operations, then model accuracy is improved, but security and privacy risks increase
Solution Approach 1:
The patent extracts and processes only the essential features and commonly accessed data patterns locally on edge devices, rather than transmitting all raw data to the cloud. This extraction approach maintains model accuracy by preserving critical information while reducing security risks by minimizing data exposure and communication requirements.
4Adaptability or versatility
If digital MAC engines are used for machine learning operations, then programmability and precision are improved, but power consumption and cost increase
Solution Approach 1:
The patent applies local quality by implementing specialized analog MAC circuits with specific optimization for common machine learning operations on edge devices. These local circuits provide sufficient programmability for inference tasks while consuming significantly less power than full digital MAC engines, matching the computational needs to the processing capability at each location.
5Speed
If large amounts of memory are integrated on the same chip as fast multipliers, then computation speed is improved, but silicon die area and cost increase
Solution Approach 1:
The patent implements partial action by integrating only the specific memory components and circuits needed for local inference operations on edge devices, rather than implementing full high-speed memory systems. This partial integration provides sufficient computation speed for common ML tasks while dramatically reducing silicon die area and cost compared to full digital MAC engine implementations.
Data Source
AI summary
Methods for performing mixed-mode Multiply-Accumulate (MAC) functions in an integrated circuit (IC) are disclosed. By performing part of the MAC operation spatially and in parallel, and part of it temporally and serially, the number of MAC operations can be programmed in the serial/temporal MAC segment as a multiple of the parallel/spatial MAC segment. Such a trait provides a degree of flexibility in programming the mixed-mode MAC function. A Programmable-Hybrid-Accumulation (PHA) method, performs the accumulation function of the MAC IC, by transforming the accumulation signal to a hybrid accumulation signal. The hybrid accumulation signal is comprised of a Most-Significant-Portion (MSP) and a Least-Significant-Portion (LSP), wherein the portions of the hybrid accumulation signal can be programmed in accordance with cost-performance objectives of an end application. Transforming the accumulated signal to a hybrid signal, and utilizing the PHA method, enables keeping the signal magnitudes bounded which prevent signal over-flow constraints while accumulation cycles proceed. Arranging a mixed-signal MAC in accordance with the PHA method can, among other benefits, help to limit the peak-to-peak analog signal swings which enhances performance attributes such as lower current consumption, faster speed, lower power supply voltage, and a wider signal accumulation range before power supply operating head-room conditions are breached.


