Voice Processor Scoring Phoneme Intervals for Data Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice processing systems face accuracy issues due to the inclusion of poor-quality voices collected from operators who may not read text correctly, leading to decreased performance in voice recognition and synthesis tasks.
Innovation Solution
A voice processing system that presents text to operators, acquires their voices, identifies phoneme output intervals, determines time length normality, and calculates a score based on phoneme correctness and occurrence frequencies, providing feedback and rewarding high-quality recordings to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If voices are collected from a large number of operators through the Internet with low cost, then the quantity of voice data increases and collection cost decreases, but the quality of voice data deteriorates due to operators failing to read text correctly
Solution Approach 1:
The system calculates a score representing correctness of voice by comparing phoneme output intervals with reference intervals, and feeds back this score information to operators. This feedback mechanism enables operators to understand their reading accuracy and improve future recordings, thereby maintaining data quality while continuing to collect from many operators.
Solution Approach 2:
The patent replaces manual quality inspection with automatic phoneme-based analysis. By converting voice to phoneme sequences and comparing with reference phoneme intervals, the system automatically identifies incorrect readings without requiring human reviewers, thus maintaining quality control while processing large volumes of data.
2Ease of operation
If operators perform recording work at their discretion, then the ease of operation increases and collection efficiency improves, but the reliability of voice data decreases due to mistakes in reading aloud
Solution Approach 1:
The system enables operators to perform self-evaluation of their recording quality through the calculated score. Operators can independently check whether their reading was correct by comparing the score against a threshold, eliminating the need for external quality verification while maintaining reliability.
Solution Approach 2:
By providing immediate score feedback after each recording, the system allows operators to self-correct and improve their reading accuracy over time, maintaining reliability while preserving operational freedom.
3Quantity of substance
If poor quality voices are included in the voice processing system, then the quantity of available voice data increases, but the accuracy of voice processing deteriorates
Solution Approach 1:
The system extracts and identifies incorrect phonemes by comparing output intervals with reference intervals, and can selectively exclude or flag these poor quality voice segments. This extraction of defective data points allows the system to maintain high accuracy by using only verified correct voice data for training and processing.
Solution Approach 2:
The patent introduces a score parameter representing voice correctness, calculated based on phoneme interval deviations. By adding this quality parameter, the system can filter and select voice data based on quality thresholds, ensuring that only high-quality voices are used for critical processing tasks while still maintaining a large overall dataset.
Data Source
AI summary
According to an embodiment, a voice processor includes a presenting unit to present text to an operator; a voice acquisition unit to acquire a voice of the operator reading aloud the text; an identifying unit to identify output intervals of phonemes included in the voice; a determination unit to determine whether each of time lengths of the output intervals is normal; a frequency acquisition unit to acquire frequency values respectively representing occurrence frequencies of contexts, respectively corresponding to the phonemes, the context including the phoneme and another phoneme adjacent to at least one side of the phoneme; and a score calculator to calculate a score representing correctness of the voice on the basis of the determination results of the time lengths of the output intervals and the frequency values of the contexts acquired respectively corresponding to the phonemes.


