Universal Phone Recognition Across Languages
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech processing applications face challenges in recognizing phones across multiple languages due to inaccuracies introduced by different phonetic transcription methods and thresholds, leading to poor performance in multi-lingual capabilities, such as voice recognition systems struggling to interpret words from different dialects like American English and British English.
Innovation Solution
A method that acquires and compares phonetic data from multiple languages, using databases, pronunciation dictionaries, and transcripts to identify matching phones in received speech, considering language validity and frequency of occurrence, to accurately recognize phones regardless of language.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple language corpora are combined for multi-lingual speech recognition, then language coverage is improved, but recognition accuracy deteriorates due to inconsistent phonetic transcription methods and thresholds across different languages
Solution Approach 1:
The patent applies homogeneity by standardizing phonetic transcription across all language corpora to a common reference standard. This ensures that phones from different languages are transcribed using consistent methods and thresholds, eliminating the inconsistency that caused accuracy deterioration when combining multiple language corpora.
Solution Approach 2:
The patent creates a universal phone set that serves all languages simultaneously. By defining a common phonetic framework that can represent phones from any language consistently, the system achieves both broad language coverage and maintained recognition accuracy across all supported languages.
2Measurement precision
If phonetic transcription uses strict matching thresholds to ensure accuracy, then transcription precision is improved, but language adaptability deteriorates because slight sound variations between dialects are misclassified
Solution Approach 1:
The patent implements dynamic threshold adjustment based on language and dialect context. Rather than using a fixed strict matching threshold, the system adapts thresholds dynamically to accommodate slight sound variations between dialects while maintaining high transcription precision for each specific language context.
Solution Approach 2:
The patent changes the transcription parameters (matching thresholds) based on the specific language and dialect being processed. This allows the system to maintain high precision for each dialect while being adaptable to their unique phonetic characteristics, resolving the contradiction between strict matching and dialectal variation.
Data Source
AI summary
Method of recognizing phones in speech of any language. Acquire phones for all languages and a set of languages. Acquire a pronunciation dictionary, a transcript of speech for the set of languages, and speech for the transcript. Receive speech containing unknown phones. If the speech's language is unknown, compare it to the phones for all languages to determine the phones. If the language is known but no phones were acquired in that language, compare the speech to the phones for all languages to determine the phones. If phones were acquired in the speech's language but no corresponding pronunciation dictionary was acquired, compare the speech to the phones for all languages to determine the phones. If a pronunciation dictionary was acquired for the phones in the speech's language but no transcript was acquired then compare the speech to the phones for all languages to determine the phones.


