Electronic Synapse Update Module for Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current neuromorphic and synaptronic systems lack efficient mechanisms for implementing spike-timing dependent plasticity (STDP) and reinforcement learning, which are crucial for mimicking biological brain functionality, particularly in the context of electronic synapses and cross-bar arrays.

Innovation Solution

The development of electronic synapses with memory elements and an update module for storing and updating states based on meta-information, utilizing a 6-terminal device with terminals for reading, setting, and resetting, and implementing STDP and reinforcement learning rules in a cross-bar array configuration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If traditional digital models are used for neuromorphic systems, then computational precision is maintained, but power consumption and device complexity increase

Engineering Contradiction:
Improvepower consumptionVSAvoiddevice complexity
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The patent replaces traditional digital electronic systems with a mechanical-inspired system using physical pendulums and gravitational forces to perform computational operations. The mechanical oscillators naturally exhibit limit cycle behavior that mimics neural spiking activity, eliminating the need for complex digital circuitry while reducing power consumption through passive mechanical energy dissipation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If STDP and reinforcement learning mechanisms are implemented in biological-like fashion, then learning capability is improved, but device complexity and manufacturing difficulty increase

Engineering Contradiction:
Improvelearning capabilityVSAvoidmanufacturing difficulty
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The mechanical oscillator system implements self-organizing learning through intrinsic physical dynamics. The pendulums automatically adjust their coupling strengths based on observed correlations between pre-synaptic and post-synaptic firing patterns, naturally implementing STDP without requiring external control circuits or complex manufacturing processes. The system serves itself by using physical interactions to encode and update synaptic weights.

Inventive Principle:
Principle #25Self-service

3Loss of time

If synchronous update mechanisms are used, then coordination is simplified, but loss of time and inability to handle delayed reinforcement signals increases

Engineering Contradiction:
Improvetime delay handlingVSAvoidcoordination complexity
Core Design Contradiction:
Loss of timeVSEase of operation

Solution Approach 1:

The system uses periodic oscillations of mechanical pendulums to naturally handle time delays in reinforcement signals. Each oscillator operates at its own natural frequency and phase, allowing asynchronous updates that preserve temporal relationships between events. The periodic nature of oscillations enables the system to accommodate variable time delays inherent in reinforcement learning scenarios while maintaining coordination through rhythmic synchronization.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentEP2641214B1Electronic synapses for reinforcement learning
Publication Date: 2017.06.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • EP2641214B1 patent drawingFigure 1A
  • EP2641214B1 patent drawingFigure 1B
  • EP2641214B1 patent drawingFigure 2

AI summary

Embodiments of the invention provide electronic synapse devices for reinforcement learning. An electronic synapse is configured for interconnecting a pre-synaptic electronic neuron and a post-synaptic electronic neuron. The electronic synapse comprises memory elements configured for storing a state of the electronic synapse and storing meta information for updating the state of the electronic synapse. The electronic synapse further comprises an update module configured for updating the state of the electronic synapse based on the meta information in response to an update signal for reinforcement learning. The update module is configured for updating the state of the electronic synapse based on the meta information, in response to a delayed update signal for reinforcement learning based on a learning rule.