Structured Weight Sparsity in Neural Network Processing Engines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Artificial neural networks (ANNs) face challenges in power efficiency and memory requirements due to their inherent structure, particularly in accessing and processing weights and activations, which are not fully optimized during the inference stage, leading to suboptimal performance and high power consumption.
Innovation Solution
The implementation of a low-power neural network architecture utilizing structured sparsity mechanisms, which identifies and leverages sparse patterns in weights and activations to reduce memory access and interlayer memory size, decoupling the control plane from the data plane and optimizing weight memory usage through guided training and synthesis methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If conventional neural network architectures are used, then computational capability is maintained, but power consumption increases and memory requirements are not optimized
Solution Approach 1:
The patent segments the neural network computations by separating weight storage and access operations from activation computations. Weight memory is organized into separate banks that can be independently accessed, allowing the system to load only necessary weight portions for current computations rather than accessing entire weight matrices, thereby reducing memory bandwidth requirements and power consumption while maintaining computational capability.
Solution Approach 2:
The patent implements preliminary weight loading and caching mechanisms where weights are pre-loaded into cache memory or buffer before being needed for computations. This allows the computational units to operate from cached data rather than continuously accessing external memory, significantly reducing memory access latency and power consumption while maintaining high computational throughput.
2Measurement precision
If full precision weights are stored and accessed, then computational accuracy is maintained, but memory requirements and data transfer volume increase
Solution Approach 1:
The patent applies different precision levels to different parts of the neural network weights based on their importance and usage patterns. Critical weights that frequently access or have high impact on computational output are stored and processed at full precision, while less critical weights can be stored at reduced precision or in compressed formats, thereby reducing overall memory requirements while maintaining computational accuracy for essential operations.
Solution Approach 2:
The patent dynamically adjusts the precision parameter of weight representations based on computational needs. The system can switch between different precision formats (e.g., full precision, reduced precision, quantized versions) of weight data depending on the specific computational task, available hardware resources, and accuracy requirements, optimizing the balance between memory usage and computational accuracy.
3Loss of energy
If structured sparsity is implemented, then memory access is reduced, but complexity of weight management increases
Solution Approach 1:
The patent introduces an intermediary weight management layer or controller that handles the complexity of structured sparsity operations. This intermediary component automatically manages weight loading, unloading, and reconfiguration based on sparsity patterns, shielding the computational units from direct complexity while enabling efficient memory access through automated weight optimization and selective loading of only non-zero weight elements.
4Ease of manufacture
If control plane and data plane are coupled, then implementation is simpler, but optimization opportunities are limited
Solution Approach 1:
The patent segments the neural network processing into distinct control plane and data plane functional blocks. The control plane handles high-level operations such as weight management, memory access control, and coordination of computational units, while the data plane performs actual computations using separated weight and activation data. This segmentation enables independent optimization of each plane, allowing the data plane to operate at high speed with minimal control interference while the control plane manages complexity centrally.
Data Source
AI summary
A novel and useful system and method of improved power performance and lowered memory requirements for an artificial neural network based on packing memory utilizing several structured sparsity mechanisms. The invention applies to neural network (NN) processing engines adapted to implement mechanisms to search for structured sparsity in weights and activations, resulting in a considerably reduced memory usage. The sparsity guided training mechanism synthesizes and generates structured sparsity weights. A compiler mechanism within a software development kit (SDK), manipulates structured weight domain sparsity to generate a sparse set of static weights for the NN. The structured sparsity static weights are loaded into the NN after compilation and utilized by both the structured weight domain sparsity mechanism and the structured activation domain sparsity mechanism. The application of structured sparsity lowers the span of search options and creates a relatively loose coupling between the data and control planes.


