Patterned Filter Clustering for DNN Memory Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional processing platforms and existing HDC encoding techniques fail to fully leverage the parallel and low-power capabilities of Hyperdimensional Computing (HDC) for edge computing applications, leading to suboptimal performance and energy efficiency, especially in deep neural networks (DNNs) with increasing size and computational demands.

Innovation Solution

An Application-Specific Integrated Circuits (ASIC) accelerator system is developed that employs patterned filter clustering and efficient encoding techniques, such as permutation encoding and XOR operations, to enhance accuracy and energy efficiency in edge computing environments, utilizing error-resilient strategies like power-gating and voltage over-scaling to reduce energy consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If Deep Neural Networks are expanded to handle increasing computational demands, then model accuracy and capability are improved, but memory requirements and computational complexity increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments filter weights into clusters and groups convolutions by clustering patterns, dividing the computational task into manageable parts. This allows the system to handle large DNN models by organizing computations through structured grouping rather than processing all weights individually, thus reducing overall computational complexity while maintaining accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges filters with identical clustering patterns into a single convolution operation, combining multiple computational tasks into one. This merging approach reduces the total number of operations required while preserving the functional equivalence of the original separate convolutions, thereby reducing computational complexity without sacrificing model accuracy.

Inventive Principle:
Principle #5Merging (Combining)

2Quantity of substance

If weight clustering is applied to compress DNN memory, then memory efficiency is improved, but computational operations must be optimized to maintain accuracy

Engineering Contradiction:
Improvememory usageVSAvoidcomputational efficiency
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent combines multiple convolution operations into a single convolution by merging filters that share identical clustering patterns. This merging reduces the total number of computational operations required while maintaining the same computational effect, thus improving productivity/computational efficiency while preserving the memory compression benefits of weight clustering.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent uses index tables to store cluster assignments for filter weights, copying the clustering information rather than storing all individual weight values. This copying approach reduces memory usage by storing only the necessary clustering indices while maintaining the ability to perform accurate computations through the index mapping.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If conventional processing platforms are used for HDC operations, then hardware availability is maintained, but parallel processing capability and energy efficiency are insufficient

Engineering Contradiction:
ImproveHDC capabilityVSAvoidenergy efficiency
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent replaces conventional CPU/GPU processing mechanisms with a specialized ASIC architecture designed for HDC operations. This substitution enables native support for hyperdimensional computing operations, including efficient handling of high-dimensional vectors and parallel operations, while significantly improving energy efficiency by eliminating the overhead of conventional processing architectures.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The ASIC architecture segments HDC operations into specialized functional units that can process high-dimensional data in parallel. This segmentation allows the system to leverage the parallel processing capability inherent in HDC while maintaining energy efficiency through dedicated hardware paths that avoid the energy overhead of general-purpose processors.

Inventive Principle:
Principle #1Segmentation

4Ease of manufacture

If existing HDC encoding techniques are used, then implementation simplicity is maintained, but prediction accuracy is insufficient for diverse applications

Engineering Contradiction:
Improveimplementation simplicityVSAvoidprediction accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent introduces dynamic dimensionality reduction capabilities that allow the system to adapt the encoding precision based on application requirements. This dynamic adjustment enables the system to maintain high prediction accuracy for diverse applications while keeping implementation manageable through configurable rather than fixed-dimensional encoding, balancing accuracy and implementation complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240428053A1Deep Neural Network Operation Via Patterned Filter Clustering and Activation Group Reuse
Publication Date: 2024.12.26 RGT UNIV OF CALIFORNIA
  • US20240428053A1 patent drawing
  • US20240428053A1 patent drawing
  • US20240428053A1 patent drawing

AI summary

Disclosed herein are techniques and architectures for enhancing the efficiency of deep neural networks through the implementation of a pattern clustering system. A pattern clustering system can enforce shared clustering topologies on filters, thereby leading to a significant reduction in memory usage through the reuse of index information. Some embodiments of the present disclosure relate to techniques for determining and assigning clustering patterns, as well as for training a network to adhere to these target patterns. Some embodiments of the present disclosure relate to an efficient accelerator based on the patterned filters. The pattern clustering system can reduce both the memory footprint and the operation count, while maintaining accuracy comparable to that of baseline models. Furthermore, the accelerator for the pattern clustering system can significantly enhance energy efficiency, surpassing the performance of conventional technologies and setting a new benchmark in the field.