Voice Labeling Error Detection via Formant Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice labeling methods, particularly those using the corpus base method for speech synthesis, face challenges in efficiently constructing voice corpora due to manual verification of labeling errors, which is labor-intensive and prone to errors, even with automatic labeling.

Innovation Solution

A voice labeling error detecting system that acquires waveform data and corresponding labeling data, classifies the data, specifies formant frequencies, calculates evaluation values, and detects deviations to automatically identify and output waveform data with labeling errors, reducing manual labor and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automatic labeling based on voice recognition is used, then productivity of voice corpus construction is improved, but labeling accuracy deteriorates due to labeling errors

Engineering Contradiction:
Improveproductivity of voice corpus constructionVSAvoidlabeling accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system performs automatic evaluation of labeling accuracy by comparing acoustic features (formant frequencies, spectral characteristics) of waveform data against the labeling data. This feedback mechanism identifies labeling errors without manual intervention, resolving the contradiction by maintaining high productivity while improving accuracy through automated verification

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system enables the voice corpus construction process to self-verify labeling accuracy through automated acoustic analysis. The evaluation unit automatically detects labeling errors by analyzing formant frequencies and spectral features, allowing the system to self-correct without external manual verification, thus maintaining both high productivity and accuracy

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If manual verification of labeling errors is performed, then labeling accuracy is improved, but productivity deteriorates due to labor-intensive process

Engineering Contradiction:
Improvelabeling accuracyVSAvoidproductivity of voice corpus construction
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system replaces the mechanical manual verification process with an automated electronic evaluation system. The evaluation unit uses computer-based acoustic analysis (formant frequency extraction, spectral analysis) to detect labeling errors, substituting human labor with automated computational methods that achieve both high accuracy and high productivity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables automatic self-verification of labeling accuracy through automated acoustic feature analysis. The evaluation unit independently assesses labeling correctness by comparing acoustic characteristics against labeling data, eliminating the need for manual verification while maintaining high productivity and accuracy

Inventive Principle:
Principle #25Self-service

3Reliability

If voice corpus contains greater number of components, then quality of speech synthesis is improved, but construction labor increases

Engineering Contradiction:
Improvequality of speech synthesisVSAvoidconstruction time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system implements automated feedback-based quality control that rapidly evaluates labeling accuracy across large numbers of voice components. This allows efficient construction of extensive voice corpora with high synthesis quality by automatically detecting and flagging labeling errors without proportionally increasing construction time

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system enables large-scale voice corpus construction with automated self-verification capabilities. The evaluation unit automatically assesses labeling accuracy across numerous components using efficient acoustic analysis, allowing high-quality speech synthesis corpora to be constructed rapidly without manual verification of each component

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS7454347B2Voice labeling error detecting system, voice labeling error detecting method and program
Publication Date: 2008.11.18 RAKUTEN GROUP INC
  • US7454347B2 patent drawing
  • US7454347B2 patent drawing
  • US7454347B2 patent drawing

AI summary

A labeling part 3 analyzes the character string data to produce a phoneme label and a prosody label, partition the voice data stored in a voice database 1 into phonemic data, and label the phonemic data, employing the phoneme label and the like. A phoneme segmenting part 4 connects the voice data labeled with the same kind of phonemic data, and a formant extracting part 5 specifies the frequency of formant of each piece of phonemic data. A processing part 6 decides an evaluation value for each phonemic data based on the frequency of formant, and an error detection part 7 detects the phonemic data of which a deviation of the evaluation value within a set of phonemic data reaches a predetermined amount.