Structured Weight Sparsity in Neural Network Processing Engines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Artificial neural networks (ANNs) face challenges in power efficiency and memory requirements due to their inherent structure, particularly in accessing and processing weights and activations, which are not fully optimized during the inference stage, leading to suboptimal performance and high power consumption.

Innovation Solution

The implementation of a low-power neural network architecture utilizing structured sparsity mechanisms, which identifies and leverages sparse patterns in weights and activations to reduce memory access and interlayer memory size, decoupling the control plane from the data plane and optimizing weight memory usage through guided training and synthesis methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If conventional neural network architectures are used, then computational capability is maintained, but power consumption increases and memory requirements are not optimized

Engineering Contradiction:
Improvepower consumptionVSAvoidcomputational capability
Core Design Contradiction:
Use of energy by moving objectVSProductivity

Solution Approach 1:

The patent segments the neural network computations by separating weight storage and access operations from activation computations. Weight memory is organized into separate banks that can be independently accessed, allowing the system to load only necessary weight portions for current computations rather than accessing entire weight matrices, thereby reducing memory bandwidth requirements and power consumption while maintaining computational capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary weight loading and caching mechanisms where weights are pre-loaded into cache memory or buffer before being needed for computations. This allows the computational units to operate from cached data rather than continuously accessing external memory, significantly reducing memory access latency and power consumption while maintaining high computational throughput.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If full precision weights are stored and accessed, then computational accuracy is maintained, but memory requirements and data transfer volume increase

Engineering Contradiction:
Improvecomputational accuracyVSAvoidmemory requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies different precision levels to different parts of the neural network weights based on their importance and usage patterns. Critical weights that frequently access or have high impact on computational output are stored and processed at full precision, while less critical weights can be stored at reduced precision or in compressed formats, thereby reducing overall memory requirements while maintaining computational accuracy for essential operations.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically adjusts the precision parameter of weight representations based on computational needs. The system can switch between different precision formats (e.g., full precision, reduced precision, quantized versions) of weight data depending on the specific computational task, available hardware resources, and accuracy requirements, optimizing the balance between memory usage and computational accuracy.

Inventive Principle:
Principle #35Parameter changes

3Loss of energy

If structured sparsity is implemented, then memory access is reduced, but complexity of weight management increases

Engineering Contradiction:
Improvememory accessVSAvoidweight management complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary weight management layer or controller that handles the complexity of structured sparsity operations. This intermediary component automatically manages weight loading, unloading, and reconfiguration based on sparsity patterns, shielding the computational units from direct complexity while enabling efficient memory access through automated weight optimization and selective loading of only non-zero weight elements.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Ease of manufacture

If control plane and data plane are coupled, then implementation is simpler, but optimization opportunities are limited

Engineering Contradiction:
Improveimplementation simplicityVSAvoidoptimization capability
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent segments the neural network processing into distinct control plane and data plane functional blocks. The control plane handles high-level operations such as weight management, memory access control, and coordination of computational units, while the data plane performs actual computations using separated weight and activation data. This segmentation enables independent optimization of each plane, allowing the data plane to operate at high speed with minimal control interference while the control plane manages complexity centrally.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11551028B2Structured weight based sparsity in an artificial neural network
Publication Date: 2023.01.10 HAILO TECH LTD
  • US11551028B2 patent drawing
  • US11551028B2 patent drawing
  • US11551028B2 patent drawing

AI summary

A novel and useful system and method of improved power performance and lowered memory requirements for an artificial neural network based on packing memory utilizing several structured sparsity mechanisms. The invention applies to neural network (NN) processing engines adapted to implement mechanisms to search for structured sparsity in weights and activations, resulting in a considerably reduced memory usage. The sparsity guided training mechanism synthesizes and generates structured sparsity weights. A compiler mechanism within a software development kit (SDK), manipulates structured weight domain sparsity to generate a sparse set of static weights for the NN. The structured sparsity static weights are loaded into the NN after compilation and utilized by both the structured weight domain sparsity mechanism and the structured activation domain sparsity mechanism. The application of structured sparsity lowers the span of search options and creates a relatively loose coupling between the data and control planes.