Acoustic Model Learning via Segmented Variation Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing acoustic model learning devices face challenges in accurately estimating affine transformation parameters due to mixed variation factors from speaker and channel differences, especially with imperfect sample data, making it difficult to separate channel and speaker variations effectively.

Innovation Solution

The proposed solution involves an acoustic model learning device with first and second variation model learning units and an environment-independent acoustic model learning unit, which estimate parameters to maximize the integrated fitness of the models to sample speech data, allowing for inverse transforms to separate and remove variations, thereby enhancing speech recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional acoustic model learning methods are used to estimate affine transformation parameters, then channel variation can be compensated, but the estimation accuracy deteriorates due to mixed speaker and channel variation factors in the data

Engineering Contradiction:
Improveaccuracy of affine transformation parameter estimationVSAvoidmixed variation factors from speaker and channel differences
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the acoustic model learning process into separate components: a speaker variation model learning unit that estimates speaker-specific parameters, a channel variation model learning unit that estimates channel-specific parameters, and an environment-independent acoustic model learning unit. This segmentation allows each unit to focus on specific variation factors independently, preventing the mixing of speaker and channel variations that plagues traditional methods.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and removes the harmful mixed variation factors by separately modeling and eliminating speaker variation components and channel variation components from the acoustic model estimation process. The environment-independent acoustic model is learned after removing these extracted variation factors, thereby obtaining accurate parameters free from contamination by speaker or channel-specific characteristics.

Inventive Principle:
Principle #2Taking out (Extraction)

2Adaptability or versatility

If sample data from multiple acoustic environments is used to improve model robustness, then adaptability improves, but measurement precision deteriorates due to difficulty in separating variation factors

Engineering Contradiction:
Improverobustness of acoustic model across different acoustic environmentsVSAvoidaccuracy of variation factor separation
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies segmentation by dividing the learning process into distinct units that handle different variation factors. The speaker variation model learning unit processes speaker-specific information, while the channel variation model learning unit processes channel-specific information. This segmented approach enables the system to utilize data from multiple acoustic environments (different speakers and channels) while maintaining precise separation of variation factors through independent parameter estimation in each unit.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8751227B2Acoustic model learning device and speech recognition device
Publication Date: 2014.06.10 NEC CORP
  • US8751227B2 patent drawing
  • US8751227B2 patent drawing
  • US8751227B2 patent drawing

AI summary

Parameters of a first variation model, a second variation model and an environment-independent acoustic model are estimated in such a way that an integrated degree of fitness obtained by integrating a degree of fitness of the first variation model to the sample speech data, a degree of fitness of the second variation model to the sample speech data, and a degree of fitness of the environment-independent acoustic model to the sample speech data becomes the maximum. Therefore, when constructing an acoustic model by using sample speech data affected by a plurality of acoustic environments; the effect on a speech which is caused by each of the acoustic environments can be extracted with high accuracy.