Spike Timing Alignment for CTC Model Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional training pipelines for GMM/HMM and DNN/HMM hybrid systems require frame-level alignment, making the training process complex and time-consuming, while end-to-end ASR systems using CTC loss function struggle with non-aligned spike timings, affecting posterior fusion and knowledge distillation between models.

Innovation Solution

A computer-implemented method for aligning spike timings of CTC models by generating a guiding model and training additional models under its guidance, using a guide loss to minimize dissimilarity in spike timing, enabling aligned posterior distributions for improved fusion and distillation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If frame-level alignment is used for training GMM/HMM and DNN/HMM systems, then posterior fusion is easy, but training process becomes complex and time-consuming

Engineering Contradiction:
Improveposterior fusion capabilityVSAvoidtraining pipeline complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the alignment requirement from the training process by using CTC loss function, which eliminates the need for frame-level alignment between input acoustic frames and output symbols. This removes the complex alignment step while preserving the ability to perform posterior fusion through spike timing alignment.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the training objective from frame-level alignment to sequence-level alignment using CTC loss. This parameter change transforms the training problem from requiring precise temporal alignment to allowing flexible alignment through the CTC algorithm, simplifying the training pipeline.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If CTC loss function is used for end-to-end ASR training, then training pipeline is simplified, but spike timing alignment between models deteriorates

Engineering Contradiction:
Improvetraining pipeline complexityVSAvoidspike timing alignment
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent introduces a guiding model as an intermediary that generates target spike timing patterns. Student models use this guiding model's spike timings as targets during training, mediating the alignment issue between CTC models without requiring complex direct alignment mechanisms.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback by using the guiding model's spike timing information to guide the training of student models. The loss function incorporates guidance from the guiding model, creating a feedback loop that aligns spike timings across models while maintaining CTC's simplified training structure.

Inventive Principle:
Principle #23Feedback

3Reliability

If multiple CTC models with arbitrary architectures are combined, then model performance improves, but computational cost increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcomputational cost
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary action by training a guiding model first to establish the target spike timing patterns. This preliminary step enables subsequent student models to be trained efficiently with aligned spike timings, reducing the need for extensive computational resources during combination and deployment.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11302309B2Aligning spike timing of models for maching learning
Publication Date: 2022.04.12 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11302309B2 patent drawing
  • US11302309B2 patent drawing
  • US11302309B2 patent drawing

AI summary

A technique for aligning spike timing of models is disclosed. A first model having a first architecture trained with a set of training samples is generated. Each training sample includes an input sequence of observations and an output sequence of symbols having different length from the input sequence. Then, one or more second models are trained with the trained first model by minimizing a guide loss jointly with a normal loss for each second model and a sequence recognition task is performed using the one or more second models. The guide loss evaluates dissimilarity in spike timing between the trained first model and each second model being trained.