Speech Data Augmentation via Vowel Frame Insertion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face accuracy degradation due to the stretching of specific speech sounds like vowel prolongations, especially in informal conversations, which are not adequately represented in traditional training data.

Innovation Solution

A computer-implemented method for data augmentation that generates partially prolonged speech data by inserting pseudo frames into original speech data, specifically at positions corresponding to vowel sounds, to simulate the stretching phenomenon and improve recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional speech data is used for training acoustic models, then the model performs well on normal utterances, but accuracy degrades on informal conversations with stretched speech sounds

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidadaptability to informal conversation
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by pre-processing the training data to insert pseudo frames that simulate vowel stretching before the acoustic model is trained. This allows the model to learn the characteristics of stretched speech sounds in advance, improving its accuracy on informal conversations without degrading performance on normal utterances.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the temporal parameter of the speech data by inserting pseudo frames at specific positions corresponding to vowel sounds. This parameter modification creates augmented training data that represents stretched speech sounds, enabling the model to adapt to informal conversation patterns while maintaining reliability on standard speech.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If speech data is augmented by inserting pseudo frames to simulate stretching, then accuracy on informal conversations improves, but data processing complexity increases

Engineering Contradiction:
Improverecognition accuracy for informal conversationVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the speech data into discrete frames and identifying specific positions corresponding to vowel sounds. Pseudo frames are inserted only at these segmented positions rather than uniformly throughout the data, which reduces processing complexity while still capturing the stretching phenomenon effectively.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses copying by creating pseudo frames that replicate the characteristics of existing speech frames and inserting them at appropriate positions. This copying approach simulates vowel stretching without requiring complex generation processes, thereby improving informal conversation recognition while keeping data processing complexity manageable.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If pseudo frames are inserted to simulate vowel stretching, then the training dataset better represents spontaneous conversation, but the original data structure is modified

Engineering Contradiction:
Improverepresentation of spontaneous conversationVSAvoiddata structure stability
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent introduces pseudo frames as intermediary elements between the original speech frames. These pseudo frames serve as mediators that simulate vowel stretching by being inserted at positions corresponding to vowel sounds, allowing the training data to better represent spontaneous conversation while maintaining a structured and organized data format.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary action by pre-calculating and marking the positions of vowel sounds in the original speech data before inserting pseudo frames. This preliminary identification of insertion positions ensures that the original data structure is systematically modified rather than randomly altered, maintaining stability while improving representation of spontaneous conversation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11227579B2Data augmentation by frame insertion for speech data
Publication Date: 2022.01.18 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11227579B2 patent drawing
  • US11227579B2 patent drawing
  • US11227579B2 patent drawing

AI summary

A technique for data augmentation for speech data is disclosed. Original speech data including a sequence of feature frames is obtained. A partially prolonged copy of the original speech data is generated by inserting one or more new frames into the sequence of the feature frames. The partially prolonged copy is output as augmented speech data for training an acoustic model for training an acoustic model.