Dynamic Neural Network Filter Masking for Learning Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural network learning methods struggle to dynamically adjust the size of filters and number of channels in hidden layers to effectively extract characteristics from input data, leading to suboptimal performance in recognizing and processing data.
Innovation Solution
A neural network learning method that utilizes a masking filter with weight information to dynamically adjust the size of specific portions of filters and the number of channels based on comparisons between output and target data, employing a gradient descent technique to minimize differences until a threshold is reached.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the size of filters and number of channels in hidden layers are fixed, then the device complexity is reduced, but the performance of the learning network model deteriorates
Solution Approach 1:
The patent applies dynamics by making the filter size and channel number adjustable during the learning process. The learning apparatus dynamically changes the size of specific portions of filters and the number of channels based on learning progress, allowing the network architecture to adapt and optimize its performance rather than remaining static throughout training.
Solution Approach 2:
The patent implements parameter changes by modifying the size of filter portions and channel numbers as learnable parameters. These parameters are updated through gradient descent based on comparison between output data and target data, enabling the system to automatically optimize architectural parameters alongside weight parameters during training.
2Measurement precision
If the size of filters and number of channels are dynamically adjusted, then the extraction of characteristics from input data is improved, but the device complexity increases
Solution Approach 1:
The patent applies dynamics by making the filter size and channel number adjustable during the learning process. The learning apparatus dynamically changes the size of specific portions of filters and the number of channels based on learning progress, allowing the network architecture to adapt and optimize its performance rather than remaining static throughout training.
Solution Approach 2:
The patent implements parameter changes by modifying the size of filter portions and channel numbers as learnable parameters. These parameters are updated through gradient descent based on comparison between output data and target data, enabling the system to automatically optimize architectural parameters alongside weight parameters during training.
3Measurement precision
If the size of specific portion in filter is increased, then the extraction of characteristics is enhanced, but the number of parameters increases leading to more training time
Solution Approach 1:
The patent applies dynamics by making the filter size and channel number adjustable during the learning process. The learning apparatus dynamically changes the size of specific portions of filters and the number of channels based on learning progress, allowing the network architecture to adapt and optimize its performance rather than remaining static throughout training.
Solution Approach 2:
The patent implements parameter changes by modifying the size of filter portions and channel numbers as learnable parameters. These parameters are updated through gradient descent based on comparison between output data and target data, enabling the system to automatically optimize architectural parameters alongside weight parameters during training.
Data Source
AI summary
The disclosure relates to an artificial intelligence (AI) system for mimicking functions, such as cognition and determination as of the human brain, by utilizing a machine learning algorithm such as deep learning, and an application thereof. Provided are a neural network learning method according to an AI system and applications thereof, the method including extracting, by using a masking filter having an effective value in a specific portion of the masking filter including weight information of at least one hidden layer included in a learning network model, characteristics of input data according to weight information of a filter corresponding to the specific portion, comparing output data with target data, the output data being obtained from the learning network model based on extracted characteristics of the input data, and updating a size of the specific portion having the effective value in the masking filter, based on a result of the comparing.


