Ternary Neural Network Weight Transformation for Edge Deployment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large and complex neural networks require significant computational power and memory, making them challenging to deploy on devices with limited resources, and existing compression techniques like network pruning and quantization often require complex post-processing and retraining, leading to performance degradation.
Innovation Solution
A method to transform pre-trained neural networks into ternary representation, allowing outputs to be determined by additions, subtractions, and bit shift operations, eliminating the need for costly Multiply-Accumulate operations without requiring retraining or complex post-processing, while maintaining performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If neural network size and complexity are increased to improve performance, then accuracy and precision are improved, but computational power and memory requirements increase
Solution Approach 1:
The patent applies parameter changes by transforming weight values from continuous floating-point parameters to discrete ternary parameters (-1, 0, +1). This fundamental parameter transformation reduces the computational complexity from multiply-accumulate operations to simpler add-subtract operations, thereby reducing energy consumption while maintaining network accuracy through the ternary decomposition approach where original weights are represented as sums of scaled ternary values
2Use of energy by moving object
If network pruning is applied to reduce computational requirements, then computational power and memory requirements are reduced, but complex post-processing and retraining are required leading to performance degradation
Solution Approach 1:
The patent applies preliminary action by performing ternary decomposition and transformation on the trained network weights before deployment. This preliminary transformation converts the network into a computationally efficient ternary form that requires no further pruning or retraining, eliminating the need for complex post-processing while maintaining the network's original accuracy through the mathematical equivalence of the ternary representation
3Use of energy by moving object
If quantization is applied to reduce precision of parameters, then computational and memory requirements are reduced, but complex post-processing is required and performance is reduced
Solution Approach 1:
The patent applies parameter changes by transforming weight values from continuous floating-point parameters to discrete ternary parameters (-1, 0, +1). This fundamental parameter transformation reduces the computational complexity from multiply-accumulate operations to simpler add-subtract operations, thereby reducing energy consumption while maintaining network accuracy through the ternary decomposition approach where original weights are represented as sums of scaled ternary values
4Productivity
If ternary transformation is applied to eliminate multiply-accumulate operations, then computational efficiency and energy efficiency are improved, but transformation complexity increases
Solution Approach 1:
The patent applies segmentation by decomposing the weight transformation process into distinct stages: (1) initial ternary quantization of weights to -1, 0, +1 values, (2) decomposition of scaled ternary values into sums of powers of two, and (3) representation as sequences of ternary vectors. This segmentation makes the transformation process more manageable and implementable through systematic processing of weight matrices layer by layer
Data Source
AI summary
The present invention relates to a computer implemented method (30) for transforming a pre-trained neural network. The method (300) comprising: receiving (S302), by a transformation device, the pre-trained neural network, wherein the pre-trained neural network comprises a number of neurons, and wherein each neuron is associated with a respective weight vector; generating (S304), by the transformation device, a ternary representation of each weight vector, by transforming each weight vector into a ternary decomposition, comprising a ternary matrix, and a power-of-two vector, wherein elements of the power-of-two vector are different powers of two; and outputting (S306), by the transformation device, a transformed neural network, wherein the weight vectors of each neuron is represented by the ternary representation; whereby an output of each neuron, obtainable by a multiplication between an input vector of each neuron and the respective weight vector, can be determined by additions, subtractions and bit shift operations.


