Neural Network Parameter Segmentation for Energy-Efficient Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The training of neural networks, particularly image classifiers, is highly CPU-intensive and energy-consuming, and existing methods like pruning lead to reduced flexibility and expressiveness, as well as potential overfitting, while deactivating processing units during inference only partially addresses energy conservation.
Innovation Solution
A method where parameters are initialized and selectively trained or retained based on a predefined criterion, such as relevance assessment or budget constraints, allowing for dynamic adjustment during training to optimize computational effort and reduce energy expenditure without completely discontinuing links between neurons, thereby improving training results and reducing overfitting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If all parameters are trained, then training accuracy is improved, but computational time and energy consumption increase
Solution Approach 1:
The patent segments the parameter set into two distinct subsets: parameters to be trained and parameters to be retained. This segmentation allows selective optimization of only critical parameters while preserving others at their initialized values, thereby reducing computational time and energy consumption while maintaining adequate training accuracy.
Solution Approach 2:
The patent applies local quality by differentiating the treatment of different parameters based on their individual importance. Critical parameters that significantly impact training accuracy are selected for optimization, while less critical parameters are retained at their initialized values. This selective approach optimizes computational resources on parameters where they yield the most benefit.
2Productivity
If parameters are pruned (set to zero), then computational effort is reduced, but flexibility and expressiveness of the ANN are reduced
Solution Approach 1:
The patent introduces dynamics by allowing flexible adjustment of the parameter subset configuration. The division between parameters to be trained and parameters to be retained can be dynamically adjusted based on training progress, performance requirements, and resource constraints. This dynamic approach enables optimization of computational efficiency while preserving the necessary flexibility and expressiveness through selective retention of non-zero parameter values.
3Use of energy by moving object
If parameters are pruned, then energy consumption is reduced, but overfitting increases
Solution Approach 1:
The patent changes the state of parameters by retaining them at their initialized non-zero values rather than setting them to zero as in traditional pruning. This parameter change maintains the network's ability to generalize while still reducing energy consumption by excluding certain parameters from the training process. The selective retention approach preserves sufficient model complexity to avoid overfitting while achieving energy efficiency.
4Use of energy by moving object
If parameters are retained instead of trained, then energy consumption is reduced, but training progress may stall
Solution Approach 1:
The patent incorporates feedback mechanisms to monitor training progress and evaluate the impact of retaining certain parameters. By assessing whether training progress stalls when parameters are retained, the system can provide feedback to adjust the configuration of the parameter subsets. This feedback loop ensures that energy-saving parameter retention decisions do not compromise overall training effectiveness.
Data Source
AI summary
A method for training an artificial neural network (ANN) whose behavior is characterized by trainable parameters. In the method, the parameters are initialized. Training data are provided which are labeled with target outputs onto which the ANN is to map the training data in each case. The training data are supplied to the ANN and mapped onto outputs by the ANN. The matching of the outputs with the learning outputs is assessed according to a predefined cost function. Based on a predefined criterion, at least one first subset of parameters to be trained and one second subset of parameters to be retained are selected from the set of parameters. The parameters to be trained are optimized. The parameters to be retained are in each case left at their initialized values or at a value already obtained during the optimization.


