Binary Activation Training for Robust Spiking Neural Networks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional deep learning neural networks are brittle to noise, including naturally occurring and adversarial attacks, and are not well-suited for implementation on reduced precision spiking or binary neuromorphic hardware.
Innovation Solution
Training neural networks with bounded ramp activation functions that converge to a discrete threshold activation, and converting them to spiking neural networks using a method that incorporates binary activations and neuromorphic hardware compatibility, while maintaining robustness and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional deep learning neural networks are used, then high computational performance can be achieved, but the system becomes brittle to noise and adversarial attacks
Solution Approach 1:
The patent transforms continuous activation functions into discrete binary activation functions by changing the parameter space. The continuous output range is mapped to discrete binary values (0 or 1), fundamentally altering the activation function's mathematical properties. This parameter transformation enables the network to achieve robustness to noise and adversarial attacks while maintaining compatibility with neuromorphic hardware.
Solution Approach 2:
The patent replaces the traditional continuous activation mechanism with a discrete binary activation mechanism. Instead of using continuous differentiable functions like ReLU or sigmoid, the system employs binary step functions that output only 0 or 1. This substitution aligns the computational model with neuromorphic hardware's natural binary spiking behavior, improving both robustness and hardware efficiency.
2Adaptability or versatility
If continuous activation functions are used in neural networks, then smooth gradient flow during training is achieved, but compatibility with reduced precision spiking and binary neuromorphic hardware is lost
Solution Approach 1:
The patent applies parameter changes by transforming the activation function's output range from continuous to discrete binary values. This parameter transformation enables direct mapping to neuromorphic hardware's binary spiking neurons, achieving hardware compatibility while maintaining sufficient precision for practical applications through the training methodology described.
Solution Approach 2:
The patent segments the continuous activation function space into discrete binary segments. By dividing the continuous output range into two distinct regions (below threshold = 0, above threshold = 1), the system creates a segmented activation function that is compatible with binary neuromorphic hardware while preserving the essential nonlinear transformation capability needed for network functionality.
3Use of energy by moving object
If binary activation functions are used, then energy efficiency on neuromorphic hardware is improved, but training convergence becomes more difficult
Solution Approach 1:
The patent applies preliminary action by performing continuous activation function training as a precursor step before transitioning to binary activation. This preliminary training phase establishes good weight initialization and network configuration, which then serves as the foundation for the subsequent binary activation training phase, thereby reducing the overall training time despite the added complexity.
Solution Approach 2:
The patent uses continuous activation functions as an intermediary during the training process. The continuous activation serves as a mediator that facilitates gradient-based optimization, allowing the network to learn effective weight configurations before the final binary activation is applied. This intermediary approach enables efficient training of binary networks by leveraging the mathematical properties of continuous functions during the learning phase.
Data Source
AI summary
A method of increasing neural network robustness. The method comprises defining an artificial neural network comprising a number of bounded ramp activation functions. The network is trained iteratively in a layer-by-layer fashion. Each iteration increases the slope of the activation functions toward a discrete threshold activation and stops when the activation functions converge to the threshold activation and the network exhibits spiking behavior. Alternatively, weight agnostic neural networks are created, wherein nodes in the networks comprise fixed shared weights. A subset of networks is identified that comprise activation functions compatible with neuromorphic hardware and are tested with a specified number of shared weight values. A score is generated for each combination of network and weight value according to performance and mapping to neuromorphic hardware, and the networks are ranked. The networks are then combined according to ranking to create a new network that exhibits spiking behavior.


