Ternary Neural Network Weight Transformation for Edge Deployment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large and complex neural networks require significant computational power and memory, making them challenging to deploy on devices with limited resources, and existing compression techniques like network pruning and quantization often require complex post-processing and retraining, leading to performance degradation.

Innovation Solution

A method to transform pre-trained neural networks into ternary representation, allowing outputs to be determined by additions, subtractions, and bit shift operations, eliminating the need for costly Multiply-Accumulate operations without requiring retraining or complex post-processing, while maintaining performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If neural network size and complexity are increased to improve performance, then accuracy and precision are improved, but computational power and memory requirements increase

Engineering Contradiction:
Improvenetwork accuracyVSAvoidcomputational power requirement
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent applies parameter changes by transforming weight values from continuous floating-point parameters to discrete ternary parameters (-1, 0, +1). This fundamental parameter transformation reduces the computational complexity from multiply-accumulate operations to simpler add-subtract operations, thereby reducing energy consumption while maintaining network accuracy through the ternary decomposition approach where original weights are represented as sums of scaled ternary values

Inventive Principle:
Principle #35Parameter changes

2Use of energy by moving object

If network pruning is applied to reduce computational requirements, then computational power and memory requirements are reduced, but complex post-processing and retraining are required leading to performance degradation

Engineering Contradiction:
Improvecomputational power requirementVSAvoidnetwork accuracy
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by performing ternary decomposition and transformation on the trained network weights before deployment. This preliminary transformation converts the network into a computationally efficient ternary form that requires no further pruning or retraining, eliminating the need for complex post-processing while maintaining the network's original accuracy through the mathematical equivalence of the ternary representation

Inventive Principle:
Principle #10Preliminary action

3Use of energy by moving object

If quantization is applied to reduce precision of parameters, then computational and memory requirements are reduced, but complex post-processing is required and performance is reduced

Engineering Contradiction:
Improvecomputational power requirementVSAvoidnetwork accuracy
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by transforming weight values from continuous floating-point parameters to discrete ternary parameters (-1, 0, +1). This fundamental parameter transformation reduces the computational complexity from multiply-accumulate operations to simpler add-subtract operations, thereby reducing energy consumption while maintaining network accuracy through the ternary decomposition approach where original weights are represented as sums of scaled ternary values

Inventive Principle:
Principle #35Parameter changes

4Productivity

If ternary transformation is applied to eliminate multiply-accumulate operations, then computational efficiency and energy efficiency are improved, but transformation complexity increases

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidtransformation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by decomposing the weight transformation process into distinct stages: (1) initial ternary quantization of weights to -1, 0, +1 values, (2) decomposition of scaled ternary values into sums of powers of two, and (3) representation as sequences of ternary vectors. This segmentation makes the transformation process more manageable and implementable through systematic processing of weight matrices layer by layer

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240046098A1Computer implemented method for transforming a pre trained neural network and a device therefor
Publication Date: 2024.02.08 INTERUNIVERSITAIR MICRO ELECTRONICS CENT (IMEC VZW)
  • US20240046098A1 patent drawing
  • US20240046098A1 patent drawing
  • US20240046098A1 patent drawing

AI summary

The present invention relates to a computer implemented method (30) for transforming a pre-trained neural network. The method (300) comprising: receiving (S302), by a transformation device, the pre-trained neural network, wherein the pre-trained neural network comprises a number of neurons, and wherein each neuron is associated with a respective weight vector; generating (S304), by the transformation device, a ternary representation of each weight vector, by transforming each weight vector into a ternary decomposition, comprising a ternary matrix, and a power-of-two vector, wherein elements of the power-of-two vector are different powers of two; and outputting (S306), by the transformation device, a transformed neural network, wherein the weight vectors of each neuron is represented by the ternary representation; whereby an output of each neuron, obtainable by a multiplication between an input vector of each neuron and the respective weight vector, can be determined by additions, subtractions and bit shift operations.