Voice Labeling Error Detection via Formant Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice labeling methods, particularly those using the corpus base method for speech synthesis, face challenges in efficiently constructing voice corpora due to manual verification of labeling errors, which is labor-intensive and prone to errors, even with automatic labeling.
Innovation Solution
A voice labeling error detecting system that acquires waveform data and corresponding labeling data, classifies the data, specifies formant frequencies, calculates evaluation values, and detects deviations to automatically identify and output waveform data with labeling errors, reducing manual labor and improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If automatic labeling based on voice recognition is used, then productivity of voice corpus construction is improved, but labeling accuracy deteriorates due to labeling errors
Solution Approach 1:
The system performs automatic evaluation of labeling accuracy by comparing acoustic features (formant frequencies, spectral characteristics) of waveform data against the labeling data. This feedback mechanism identifies labeling errors without manual intervention, resolving the contradiction by maintaining high productivity while improving accuracy through automated verification
Solution Approach 2:
The system enables the voice corpus construction process to self-verify labeling accuracy through automated acoustic analysis. The evaluation unit automatically detects labeling errors by analyzing formant frequencies and spectral features, allowing the system to self-correct without external manual verification, thus maintaining both high productivity and accuracy
2Manufacturing precision
If manual verification of labeling errors is performed, then labeling accuracy is improved, but productivity deteriorates due to labor-intensive process
Solution Approach 1:
The system replaces the mechanical manual verification process with an automated electronic evaluation system. The evaluation unit uses computer-based acoustic analysis (formant frequency extraction, spectral analysis) to detect labeling errors, substituting human labor with automated computational methods that achieve both high accuracy and high productivity
Solution Approach 2:
The system enables automatic self-verification of labeling accuracy through automated acoustic feature analysis. The evaluation unit independently assesses labeling correctness by comparing acoustic characteristics against labeling data, eliminating the need for manual verification while maintaining high productivity and accuracy
3Reliability
If voice corpus contains greater number of components, then quality of speech synthesis is improved, but construction labor increases
Solution Approach 1:
The system implements automated feedback-based quality control that rapidly evaluates labeling accuracy across large numbers of voice components. This allows efficient construction of extensive voice corpora with high synthesis quality by automatically detecting and flagging labeling errors without proportionally increasing construction time
Solution Approach 2:
The system enables large-scale voice corpus construction with automated self-verification capabilities. The evaluation unit automatically assesses labeling accuracy across numerous components using efficient acoustic analysis, allowing high-quality speech synthesis corpora to be constructed rapidly without manual verification of each component
Data Source
AI summary
A labeling part 3 analyzes the character string data to produce a phoneme label and a prosody label, partition the voice data stored in a voice database 1 into phonemic data, and label the phonemic data, employing the phoneme label and the like. A phoneme segmenting part 4 connects the voice data labeled with the same kind of phonemic data, and a formant extracting part 5 specifies the frequency of formant of each piece of phonemic data. A processing part 6 decides an evaluation value for each phonemic data based on the frequency of formant, and an error detection part 7 detects the phonemic data of which a deviation of the evaluation value within a set of phonemic data reaches a predetermined amount.


