Phonetic Labeling Consistency Across Diverse Speech Corpora
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech processing applications face challenges in combining corpora with phonetic information generated by different methods or people, leading to inconsistencies and poor performance when processing multiple languages, as seen in voice recognition systems struggling to interpret variations in pronunciation between languages like American English and British English.
Innovation Solution
A method to combine corpora by generating consistent phonetic information, involving the selection of corpora, generation of phonetic transcripts using user-definable dictionaries, identification of allophones, and replacement of phone symbols to align phonetic labeling across different corpora, ensuring consistency and minimizing inaccuracies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If corpora are combined without phonetic consistency processing, then the quantity of speech data increases, but phonetic labeling consistency deteriorates
Solution Approach 1:
The patent transforms phonetic labels from different transcription standards into a unified parameter space by mapping various phonetic notations (IPA, ARPABET, etc.) to a common set of phone symbols. This parameter transformation enables consistent processing of multi-lingual corpora while preserving the original phonetic information through reversible mapping relationships.
Solution Approach 2:
The patent introduces an intermediary phonetic alignment layer that acts as a mediator between different phonetic transcription systems. This intermediary layer uses pronunciation dictionaries and alignment algorithms to translate between different phonetic notations, enabling consistent combination of corpora from multiple sources without direct conflict between incompatible labeling systems.
2Ease of manufacture
If different phonetic transcription methods are used for different corpora, then phonetic information can be generated for each corpus independently, but inconsistencies are introduced when combining corpora
Solution Approach 1:
The patent creates a universal phonetic labeling framework that can handle multiple transcription standards simultaneously. The system maintains compatibility with different phonetic notations while providing a unified output format, allowing independent corpus generation to continue using existing methods while ensuring consistency when corpora are combined through the universal mapping layer.
Solution Approach 2:
The patent applies parameter transformation to convert phonetic labels from various transcription methods into a standardized parameter set. By changing the parameter representation of phonetic symbols rather than the underlying phonetic content, the system preserves the independence of corpus generation while achieving consistency in the combined corpus.
3Productivity
If phonetic transcripts are generated using forced alignment with different thresholds, then each corpus can be processed efficiently, but accuracy deteriorates when combining corpora with different threshold criteria
Solution Approach 1:
The patent performs preliminary phonetic alignment and threshold standardization during the corpus processing stage rather than during combination. By pre-processing each corpus with its own optimal threshold while maintaining a mapping to the unified phonetic system, the system preserves processing efficiency while ensuring accuracy in the combined corpus through the preliminary alignment work.
Data Source
AI summary
The present invention is a method of combining corpora to achieve consistency in phonetic labeling. Corpora are received. A first corpus is selected from the corpora. Generating a phonetic transcript if the first corpus does not include one. A second corpus is selected from the corpora. Generating a phonetic transcript if the second corpus does not include one. Each allophone in the second corpus is identified. At least one allophone is identified for each phone in the second corpus. For each phone in the second corpus, the allophone to which it most closely matches is identified. Each phone symbol in the phone transcript of the second corpus is replaced with a symbol for the corresponding identified allophone. The first corpus and second corpus are combined, including their phonetic transcripts, and designated as the first corpus. If there is another corpus in the corpora to be processed return to the step of selecting another second corpus.

