Optimum-Partitioned Neural Network for Phoneme Boundary Tuning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automatic phoneme labeling techniques, such as those using Hidden Markov Models (HMM) and neural networks, face challenges in accurately reflecting acoustic changes at phoneme boundaries, leading to high errors and inefficiencies in speech recognition and synthesis processes.
Innovation Solution
The method employs an optimum-partitioned classified neural network of the MLP type to segment phoneme combinations into partitions, searching for neural networks with minimum errors and updating weights to converge on a total error sum, thereby tuning phoneme boundaries for improved accuracy and speed in automatic labeling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If HMM probability modeling is used for phoneme segmentation, then a generated model for total training data can be considered as an optimum model, but it cannot reflect physical characteristics related to acoustic feature variables of a speech signal
Solution Approach 1:
The patent segments the speech signal into distinct phoneme units by detecting phoneme boundaries, separating the probabilistic modeling aspect (HMM for overall structure) from the acoustic characteristic aspect (time-domain features for boundary detection). This allows each method to operate in its optimal domain while combining their strengths.
Solution Approach 2:
The patent introduces an intermediary phoneme boundary detection mechanism that bridges HMM probability modeling and acoustic feature analysis. The boundary detector uses time-domain acoustic features as an intermediary to translate probabilistic model outputs into physically meaningful phoneme boundaries that reflect actual acoustic characteristics.
2Measurement precision
If speech segmentation techniques reflecting acoustic changes are used, then acoustic feature variable transitions can be segmented, but it is difficult to directly adapt to automatic labeling due to limited context information consideration
Solution Approach 1:
The patent merges speech segmentation techniques with automatic labeling by integrating phoneme boundary detection results directly into the labeling process. The segmented phoneme boundaries from acoustic feature analysis are combined with HMM-based probability modeling to create a unified automatic labeling system that maintains both acoustic precision and labeling versatility.
Solution Approach 2:
The patent changes the parameter representation by transforming acoustic feature variables into phoneme boundary indicators that can be directly used for labeling. This parameter transformation allows acoustic change detection to be seamlessly adapted to automatic labeling requirements while preserving the physical characteristics of speech signals.
3Ease of operation
If neural networks are used to detect phoneme boundaries, then automatic labeling can be performed without probability modeling, but learned coefficients frequently converge to local optimum and not global optimum
Solution Approach 1:
The patent applies preliminary phoneme boundary detection using time-domain acoustic features before applying neural network-based automatic labeling. This preliminary action provides a reliable initial framework that guides the neural network optimization process, preventing convergence to poor local optima and improving overall labeling accuracy.
Solution Approach 2:
The patent implements feedback mechanisms where phoneme boundary detection results are used to refine neural network coefficient learning. The system continuously adjusts learned coefficients based on feedback from acoustic feature analysis, ensuring convergence toward global optima rather than getting trapped in local optima, thereby improving labeling reliability.
Data Source
AI summary
A method of automatic labeling using an optimum-partitioned classified neural network includes searching for neural networks having minimum errors with respect to a number of L phoneme combinations from a number of K neural network combinations generated at an initial stage or updated, updating weights during learning of the K neural networks by K phoneme combination groups searched with the same neural networks, and composing an optimum-partitioned classified neural network combination using the K neural networks of which a total error sum has converged; and tuning a phoneme boundary of a first label file by using the phoneme combination group classification result and the optimum-partitioned classified neural network combination, and generating a final label file reflecting the tuning result.


