Neural Network Weight Pruning for Memory and Accuracy Trade-off
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Artificial neural networks face challenges in miniaturization and commercialization due to increased complexity and memory usage, leading to overfitting issues and decreased reliability in predicting new data, necessitating a method to reduce system cost while maintaining performance.
Innovation Solution
A method and apparatus for compressing artificial neural networks by adjusting weights among layers, determining a compression rate based on initial and compressed task accuracy, and re-compressing the network to optimize performance and reduce complexity, involving a controller that evaluates and adjusts weights to minimize accuracy loss.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the amount of learning of an artificial neural network increases, then the accuracy of previously learned data increases, but the reliability of prediction values regarding new data decreases (overfitting problem)
Solution Approach 1:
The patent extracts and removes redundant connections (pruning) from the neural network to reduce overfitting. By identifying and eliminating unnecessary weight connections, the network maintains accuracy on training data while improving generalization to new data, directly addressing the overfitting problem described in the contradiction.
Solution Approach 2:
The patent changes the parameter of connection density by compressing the neural network. Through adaptive compression techniques, the network adjusts its complexity level, reducing the number of active connections while maintaining performance, thereby transforming the trade-off between memorization and generalization.
2Measurement precision
If the complexity of the artificial neural network increases, then the accuracy of previously learned data increases, but memory use increases causing problems in miniaturization and commercialization
Solution Approach 1:
The patent extracts and removes redundant connections (pruning) from the neural network to reduce memory usage. By identifying and eliminating unnecessary weight connections, the network maintains accuracy on training data while reducing the computational resources required, directly addressing the memory consumption issue described in the contradiction.
Solution Approach 2:
The patent changes the parameter of connection density by compressing the neural network. Through adaptive compression techniques, the network adjusts its complexity level, reducing the number of active connections while maintaining performance, thereby transforming the trade-off between accuracy and memory usage.
3Quantity of substance
If a compression technique is applied to reduce system cost, then memory use decreases, but the performance of the artificial neural network may deteriorate
Solution Approach 1:
The patent implements feedback mechanisms during compression to monitor and maintain network performance. By continuously evaluating accuracy metrics and adjusting compression parameters accordingly, the system ensures that performance degradation is minimized while achieving memory reduction goals, directly addressing the performance deterioration risk.
Solution Approach 2:
The patent changes the parameter of compression intensity adaptively. Through multi-stage compression with varying compression rates, the system finds the optimal balance point where memory usage is reduced but performance remains acceptable, transforming the binary trade-off between compression and performance.
Data Source
AI summary
Provided are an apparatus and method of compressing an artificial neural network. According to the method and the apparatus, an optimal compression rate and an optimal operation accuracy are determined by compressing an artificial neural network, determining a task accuracy of a compressed artificial neural network, and automatically calculating a compression rate and a compression ratio based on the determined task accuracy. The method includes obtaining an initial value of a task accuracy for a task processed by the artificial neural network, compressing the artificial neural network by adjusting weights of connections among layers of the artificial neural network included in information regarding the connections, determining a compression rate for the compressed artificial neural network based on the initial value of the task accuracy and a task accuracy of the compressed artificial neural network, and re-compressing the compressed artificial neural network according to the compression rate.


