Speech Recognition Model Data Augmentation via Contextual Acoustic Parameter Changes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition models face challenges in efficiently augmenting learning data, leading to increased augmentation time and potential reduction in performance due to inclusion of unnecessary data, which can result in overfitting and decreased model performance.
Innovation Solution
An electronic device performs natural language understanding on learning data to obtain information about the context of speech utterance, including acoustic features, and generates additional speech signals based on this information to augment the learning data, thereby training the speech recognition model with context-specific data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If learning data is augmented in a random manner to avoid overfitting, then the diversity of learning data is improved, but unnecessary data may be included which increases augmentation time and reduces model performance
Solution Approach 1:
The patent changes the parameters of speech signals systematically based on acoustic features extracted from natural language understanding, rather than applying random transformations. This involves modifying pitch, speed, and other acoustic parameters according to the semantic context and utterance situation, thereby generating diverse yet relevant augmented data without unnecessary random variations.
Solution Approach 2:
The patent replaces the mechanical random transformation approach with an intelligent system that uses natural language understanding and acoustic feature analysis. Instead of blindly applying random augmentations, the system substitutes a sophisticated processing mechanism that understands the semantic meaning and generates context-appropriate variations, reducing wasted computations.
2Reliability
If learning data is augmented in a random manner, then overfitting is avoided, but model performance may be reduced due to inclusion of unnecessary data
Solution Approach 1:
The system systematically varies acoustic parameters (pitch, speed, tone) based on the semantic context and utterance situation derived from natural language understanding. This creates diverse training samples that reflect real-world variations while maintaining relevance to the target task, avoiding both overfitting and inclusion of irrelevant data.
Solution Approach 2:
The patent applies different augmentation strategies to different parts of the speech signal based on local acoustic characteristics and contextual information. Rather than uniform random transformation, the system tailors the augmentation to specific acoustic features and semantic contexts, ensuring each augmented sample maintains appropriate quality and relevance.
3Manufacturing precision
If a large amount of learning data is constructed to improve speech recognition model performance, then model accuracy is improved, but costs and time required to build the data increase significantly
Solution Approach 1:
The patent creates multiple copies and variations of existing speech data through systematic parameter transformation based on acoustic features. Instead of collecting large amounts of new data, the system generates diverse training samples by copying and transforming existing high-quality data, significantly reducing data construction time while maintaining model accuracy.
Solution Approach 2:
The system performs preliminary natural language understanding and acoustic feature extraction on the original data to prepare transformation parameters in advance. This preliminary analysis enables efficient generation of augmented data without requiring time-consuming manual annotation or collection processes for each new sample.
Data Source
AI summary
Disclosed are an electronic device and a method of controlling the electronic device. An electronic device according to an embodiment may perform a method comprising: performing natural language understanding for a first text included in learning data, obtaining first information associated with a speech corresponding to the first text being uttered based on a result of the natural language understanding, obtain second information associated with an acoustic feature corresponding to the speech corresponding to the first text being uttered based on the first information, obtaining a plurality of speech signals corresponding to the first text by converting a first speech signal corresponding to the first text based on the first information and the second information, and training a speech recognition model based on the plurality of obtained speech signals and the first text.


