Phoneme Boundary Labeling Error Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic phoneme labeling methods using deep neural networks often result in phoneme boundaries that differ significantly from manually applied labels, leading to errors in speech synthesis, requiring costly manual corrections and extensive listening time for verification.
Innovation Solution
A labeling processing method that includes forward and backward labeling steps to generate label information and a learning step to detect appropriate phoneme boundaries based on the difference between time information in both directions, using a model to determine labeling errors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automatic labeling using deep neural networks is used, then labeling speed and productivity are improved, but labeling precision deteriorates
Solution Approach 1:
The patent applies backward labeling in addition to forward labeling. The backward labeling step processes the speech signal from end to beginning, creating a second set of time information. By comparing the time information from forward and backward labeling, the system can detect inconsistencies and correct phoneme boundary positions, thereby improving precision while maintaining automated processing.
Solution Approach 2:
The system uses the difference between forward and backward labeling results as feedback to detect and correct labeling errors. The learning model processes this feedback information to identify inappropriate phoneme boundaries and generate corrected labels, creating a self-correcting automated labeling system that improves precision without sacrificing productivity.
2Measurement precision
If manual correction of all phoneme boundaries is performed, then labeling precision is improved, but loss of time and cost increase
Solution Approach 1:
The system performs self-correction by automatically detecting labeling errors through the comparison of forward and backward labeling results. The learning model processes the difference information and generates corrected phoneme boundary positions without requiring manual intervention, enabling the system to correct its own errors while maintaining high precision.
Solution Approach 2:
The patent replaces the mechanical manual correction process with an automated learning-based correction system. Instead of requiring human annotators to review and correct each phoneme boundary, the system uses computational models that process the labeling data and automatically generate corrections, eliminating the time-consuming manual correction step while maintaining or improving precision.
3Reliability
If manual verification of all labeling targets is performed, then reliability is improved, but loss of time increases
Solution Approach 1:
Instead of requiring complete manual verification of all labeling targets, the system performs partial verification by focusing on detecting labeling errors through the forward-backward labeling comparison. The learning model identifies and corrects only the problematic boundaries, allowing rapid verification of reliability without the need to manually review every single phoneme boundary in detail.
Data Source
AI summary
A labeling processing device (100) generates first label information by labeling time information in a forward direction with respect to a plurality of phoneme boundaries set in speech information for learning. The labeling processing device (100) generates second label information by labeling time information in a direction opposite to the forward direction with respect to a plurality of phoneme boundaries set in speech information for learning and inverting the order of the labeled time information. The labeling processing device (100) detects whether phoneme boundaries are appropriate on the basis of a difference between time information on a plurality of phoneme boundaries included in the first label information and time information on a plurality of phoneme boundaries included in the second label information.


