Quaternion Deep Neural Network Gradient Computation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional neural networks based on real-number calculus struggle to effectively incorporate quaternions in deep multi-layered neural networks, leading to pseudo-gradients and failure to satisfy standard calculus rules, limiting their application in machine learning tasks that require quaternion operations.
Innovation Solution
Adapting machine learning systems to use quaternion-specific computations, where each neuron stores input, output, weighting, and bias values as n-dimensional quaternions, enabling consistent quaternion operations for gradient-based training and backpropagation, leveraging quaternion properties like rotational invariance for improved performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If conventional coordinate-wise real-number calculus is applied to quaternions, then the computational operations can be performed, but the standard product or chain rules of calculus are not satisfied and pseudo-gradients are generated
Solution Approach 1:
The patent changes the fundamental parameter of calculus from real-number based to quaternion-based by defining new quaternion differential operators and gradient definitions that respect quaternion algebra properties. This includes defining left and right quaternion derivatives, quaternion chain rules, and quaternion product rules that maintain mathematical consistency within the quaternion domain rather than applying real-number calculus coordinate-wise.
2Reliability
If quaternions are incorporated into deep multi-layered neural networks, then rotational invariance and singularity-free representations are achieved, but the training of hidden layers becomes difficult due to lack of proper quaternion calculus rules
Solution Approach 1:
The patent substitutes the mechanical coordinate-wise differentiation approach with a unified quaternion calculus system. By defining quaternion-specific differentiation operators, chain rules, and backpropagation algorithms that operate directly on quaternion-valued neurons, the system replaces the complex coordination of multiple real-valued calculations with a more elegant quaternion-native approach that simplifies the overall training mechanism.
Solution Approach 2:
The patent segments the quaternion calculus into distinct components: quaternion addition, quaternion multiplication, quaternion transpose, quaternion conjugate, left quaternion derivative, right quaternion derivative, and quaternion chain rules. Each component is defined separately with clear mathematical properties, allowing systematic application in neural network training while maintaining overall quaternion consistency.
3Productivity
If quaternion operations are used in neural networks, then parameter reduction and computational efficiency are achieved, but conventional training methods fail to properly train hidden layers
Solution Approach 1:
The patent enables quaternion-based neural networks to be self-sufficient by defining complete quaternion calculus rules including derivatives, chain rules, and backpropagation algorithms that operate entirely within the quaternion domain. This eliminates the need to fall back on real-number calculus approximations and allows the quaternion network to train itself using native quaternion operations, maintaining both efficiency and mathematical consistency.
Data Source
AI summary
A quaternion deep neural network (QTDNN) includes a plurality of modular hidden layers, each comprising a set of QT computation sublayers, including a quaternion (QT) general matrix multiplication sublayer, a QT non-linear activations sublayer, and a QT sampling sublayer arranged along a forward signal propagation path. Each QT computation sublayer of the set has a plurality of QT computation engines. In each modular hidden layer, a steering sublayer precedes each of the QT computation sublayers along the forward signal propagation path. The steering sublayer directs a forward-propagating quaternion-valued signal to a selected at least one QT computation engine of a next QT computation subsequent sublayer.


