Highway Network Gating Mechanism for Deep Neural Network Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Training deeper neural networks is challenging due to optimization difficulties, especially as the number of layers increases, leading to issues with computational and statistical efficiency in complex tasks.
Innovation Solution
The introduction of a learned gating mechanism in neural networks, referred to as highway networks, which allows for information to flow across multiple layers without attenuation, enabling the optimization of networks with virtually arbitrary depth using simple Stochastic Gradient Descent (SGD) with momentum.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the number of layers in a neural network is increased to improve representation efficiency and computational performance, then the network depth increases, but optimization difficulty increases significantly making training challenging
Solution Approach 1:
The patent introduces skip connections as intermediary pathways that directly connect input layers to output layers, bypassing intermediate transformation layers. This allows gradients to flow more effectively through the network during backpropagation, solving the optimization difficulty in deep networks while maintaining the benefits of increased network depth for computational efficiency
2Device complexity
If traditional neural network architectures are used, then simple structures are maintained, but information flow is attenuated across multiple layers limiting depth
Solution Approach 1:
The patent segments the information flow into two distinct pathways: direct skip connections that preserve original information without transformation, and transformation layers that process information through conventional neural network operations. This segmentation allows information to flow across multiple layers without attenuation while maintaining structural simplicity
Data Source
AI summary
A computer-based method includes receiving an input signal at a neuron in a computer-based neural network that includes a plurality of neuron layers, applying a first non-linear transform to the input signal at the neuron to produce a plain signal, and calculating a weighted sum of a first component of the input signal and the plain signal at the neuron. In a typical implementation, the first non-linear transform is a function of the first component of the input signal and at least a second component of the input signal.


