Asymmetric RPU DNN Training via Matrix Decomposition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep neural networks (DNNs) trained with asymmetric resistive processing unit (RPU) devices face challenges due to non-linear and non-symmetric switching characteristics, which hinder accurate weight adjustments and impair the implementation of backpropagation and stochastic gradient descent (SGD) during training.

Innovation Solution

The method involves using a weight matrix as a linear combination of two matrices, A and C, where matrix A stores conductance values in RPU devices and matrix C stores zero-weight conductance values, allowing for symmetric weight updates around a zero-point, thereby tolerating hardware bias and improving training accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If asymmetric RPU devices are used for DNN training, then device complexity is reduced and manufacturing is simplified, but weight update accuracy deteriorates due to non-symmetric switching characteristics

Engineering Contradiction:
ImproveRPU device fabricationVSAvoidweight update accuracy
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent explicitly embraces the asymmetric nature of RPU devices rather than attempting to correct it. By designing the training algorithm to work with asymmetric devices from the outset, the system accepts the non-symmetric switching characteristics as a fundamental property and develops compensation strategies that leverage this reality, thereby resolving the contradiction between ease of manufacture and weight update accuracy

Inventive Principle:
Principle #4Asymmetry

Solution Approach 2:

The patent modifies the training parameters and algorithms to accommodate asymmetric RPU behavior. By changing the approach to weight updates and introducing compensation mechanisms that account for directional biases in conductance changes, the system maintains training accuracy while using simpler asymmetric devices

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If standard backpropagation is used with asymmetric RPUs, then training process simplicity is maintained, but training accuracy deteriorates due to hardware bias

Engineering Contradiction:
Improvetraining process simplicityVSAvoidtraining accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces feedback mechanisms that monitor the asymmetric behavior of RPU devices during training and adjust subsequent weight updates accordingly. By measuring the actual conductance changes and comparing them to expected values, the system generates corrective feedback that compensates for hardware biases, thereby maintaining training accuracy without significantly complicating the training process

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies preliminary calibration and characterization of RPU devices before training begins. By pre-characterizing the asymmetric switching behavior of each device and storing compensation factors, the system prepares corrective measures in advance that can be applied during training, thus maintaining simplicity while improving accuracy

Inventive Principle:
Principle #10Preliminary action

3Stability of the object's composition

If symmetric weight updates are applied to asymmetric RPUs, then theoretical training correctness is maintained, but actual weight adjustment precision deteriorates due to non-linear switching

Engineering Contradiction:
Improvetraining algorithm correctnessVSAvoidweight adjustment precision
Core Design Contradiction:
Stability of the object's compositionVSManufacturing precision

Solution Approach 1:

The patent transitions from static, fixed weight update rules to dynamic update mechanisms that adapt to the actual state of RPU devices. By making the update process dynamic and state-dependent, the system can adjust its approach based on real-time device behavior, thereby maintaining both theoretical correctness and practical precision despite asymmetric non-linear switching characteristics

Inventive Principle:
Principle #15Dynamics

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables superior training results by minimizing the objective function and internal energy of RPU devices, providing more accurate and stable weight updates despite non-ideal RPU characteristics.

Implementation Method 1

each RPU includes a first terminal, a second terminal and an active region. A conductance state of the active region identifies a weight value of the RPU, which can be updated/adjusted by application of a signal to the first/second terminals

Methodology Applied
Scientific EffectResistive switching: Electrical Resistance

Implementation Method 2

transmitting, in a forward cycle, an input vector x as voltage pulses through the conductive column wires of the cross-point array A and the cross-point array C and reading a resulting output vector y as current output from the conductive row wires

Methodology Applied
Scientific EffectOhm's law: Ohm's Law

Data Source

PatentUS11562249B2DNN training with asymmetric RPU devices
Publication Date: 2023.01.24 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11562249B2 patent drawing
  • US11562249B2 patent drawing
  • US11562249B2 patent drawing

AI summary

In a method of training a DNN, a weight matrix (W) is provided as a linear combination of matrices/arrays A and C. In a forward cycle, an input vector x is transmitted through arrays A and C and output vector y is read. In a backward cycle, an error signal δ is transmitted through arrays A and C and output vector z is read. Array A is updated by transmitting input vector x and error signal δ through array A. In a forward cycle, an input vector ei is transmitted through array A and output vector y′ is read. ƒ(y′) is calculated using y′. Array C is updated by transmitting input vector ei and ƒ(y′) through array C. A DNN is also provided.