Coefficient Mapping for Matrix Multiply Accelerator Utilization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional artificial intelligence systems, particularly neural network models, require significant energy and compute power due to their large circuitry, leading to latency issues when deployed in remote edge devices, making them bulky and unsustainable.
Innovation Solution
Configuring an array of matrix multiply accelerators with coefficient mapping techniques to optimize computational utilization, including input/output handling, and partitioning regions based on application requirements, allowing for efficient deployment in edge devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If traditional computers and systems are used to implement neural network models, then compute power and data processing speed are improved, but device size and energy consumption increase significantly
Solution Approach 1:
The patent segments the neural network computation into distinct phases (training and inference) and implements specialized hardware architectures for each phase. The training system uses traditional high-power compute resources, while the inference system employs energy-efficient dedicated hardware accelerators that can be deployed at the edge, thus resolving the contradiction between compute power and energy consumption.
Solution Approach 2:
The patent introduces an intermediary component - a specialized inference engine or accelerator - that bridges the gap between the training phase on traditional systems and the deployment phase at the edge. This intermediary hardware is optimized for inference workloads, providing sufficient compute power while consuming significantly less energy than general-purpose computers.
2Power
If traditional computers with large circuitry are deployed at edge devices, then compute power is improved, but device size becomes bulky and impractical
Solution Approach 1:
The patent applies local quality by implementing heterogeneous computing architectures where different components have specialized functions. Instead of using uniform high-power processors throughout, the system employs dedicated inference accelerators with optimized circuitry specifically for neural network inference, reducing overall device size while maintaining necessary compute power at the edge.
Solution Approach 2:
The patent employs dynamic voltage and frequency scaling (DVFS) and adaptive resource allocation in the edge deployment. The inference engine can dynamically adjust its power consumption and computational capacity based on workload demands, allowing the device to maintain adequate compute power when needed while minimizing size and power consumption during idle or low-demand periods.
3Power
If remote computing systems are used for neural network inference, then compute power is sufficient, but latency increases due to network transmission delays
Solution Approach 1:
The patent implements preliminary action by pre-training neural network models on centralized systems and then deploying the trained model weights to edge devices in advance. This allows the edge devices to perform inference locally without needing to transmit data back and forth with remote systems, eliminating network latency while maintaining sufficient compute power through the pre-deployed models.
Solution Approach 2:
The patent extracts the inference capability from remote computing systems and places it directly at the edge devices. By taking out the neural network model and deploying it locally, the system eliminates the need for continuous network communication during inference, thus resolving the latency issue while maintaining adequate compute power through specialized local hardware.
Data Source
AI summary
Systems and methods of configuring a fixed memory array of an integrated circuit with coefficients of one or more applications includes identifying a utilization constraint type of the fixed memory array from a plurality of distinct utilization constraint types based on computing attributes of the one or more applications; identifying at least one coefficient mapping technique from a plurality of distinct coefficient mapping techniques that addresses the utilization constraint type; configuring the fixed memory array according to the at least one coefficient mapping technique, wherein configuring the array includes at least setting within the array the coefficients of the one or more applications in an arrangement prescribed by the at least one coefficient mapping technique that optimizes a computational utilization of the fixed memory array.


