Optimum-Partitioned Neural Network for Phoneme Boundary Tuning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automatic phoneme labeling techniques, such as those using Hidden Markov Models (HMM) and neural networks, face challenges in accurately reflecting acoustic changes at phoneme boundaries, leading to high errors and inefficiencies in speech recognition and synthesis processes.

Innovation Solution

The method employs an optimum-partitioned classified neural network of the MLP type to segment phoneme combinations into partitions, searching for neural networks with minimum errors and updating weights to converge on a total error sum, thereby tuning phoneme boundaries for improved accuracy and speed in automatic labeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If HMM probability modeling is used for phoneme segmentation, then a generated model for total training data can be considered as an optimum model, but it cannot reflect physical characteristics related to acoustic feature variables of a speech signal

Engineering Contradiction:
Improvemodel optimalityVSAvoidacoustic characteristic reflection
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent segments the speech signal into distinct phoneme units by detecting phoneme boundaries, separating the probabilistic modeling aspect (HMM for overall structure) from the acoustic characteristic aspect (time-domain features for boundary detection). This allows each method to operate in its optimal domain while combining their strengths.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary phoneme boundary detection mechanism that bridges HMM probability modeling and acoustic feature analysis. The boundary detector uses time-domain acoustic features as an intermediary to translate probabilistic model outputs into physically meaningful phoneme boundaries that reflect actual acoustic characteristics.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If speech segmentation techniques reflecting acoustic changes are used, then acoustic feature variable transitions can be segmented, but it is difficult to directly adapt to automatic labeling due to limited context information consideration

Engineering Contradiction:
Improveacoustic change segmentationVSAvoidautomatic labeling adaptation
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent merges speech segmentation techniques with automatic labeling by integrating phoneme boundary detection results directly into the labeling process. The segmented phoneme boundaries from acoustic feature analysis are combined with HMM-based probability modeling to create a unified automatic labeling system that maintains both acoustic precision and labeling versatility.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent changes the parameter representation by transforming acoustic feature variables into phoneme boundary indicators that can be directly used for labeling. This parameter transformation allows acoustic change detection to be seamlessly adapted to automatic labeling requirements while preserving the physical characteristics of speech signals.

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If neural networks are used to detect phoneme boundaries, then automatic labeling can be performed without probability modeling, but learned coefficients frequently converge to local optimum and not global optimum

Engineering Contradiction:
Improveautomatic labeling capabilityVSAvoidlabeling accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent applies preliminary phoneme boundary detection using time-domain acoustic features before applying neural network-based automatic labeling. This preliminary action provides a reliable initial framework that guides the neural network optimization process, preventing convergence to poor local optima and improving overall labeling accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where phoneme boundary detection results are used to refine neural network coefficient learning. The system continuously adjusts learned coefficients based on feedback from acoustic feature analysis, ensuring convergence toward global optima rather than getting trapped in local optima, thereby improving labeling reliability.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS7444282B2Method of setting optimum-partitioned classified neural network and method and apparatus for automatic labeling using optimum-partitioned classified neural network
Publication Date: 2008.10.28 SAMSUNG ELECTRONICS CO LTD
  • US7444282B2 patent drawing
  • US7444282B2 patent drawing
  • US7444282B2 patent drawing

AI summary

A method of automatic labeling using an optimum-partitioned classified neural network includes searching for neural networks having minimum errors with respect to a number of L phoneme combinations from a number of K neural network combinations generated at an initial stage or updated, updating weights during learning of the K neural networks by K phoneme combination groups searched with the same neural networks, and composing an optimum-partitioned classified neural network combination using the K neural networks of which a total error sum has converged; and tuning a phoneme boundary of a first label file by using the phoneme combination group classification result and the optimum-partitioned classified neural network combination, and generating a final label file reflecting the tuning result.