Language Identification Combining Acoustic and Textual Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current language identification techniques, such as acoustic language identification and textual verification, suffer from high error rates and resource-intensive processing, making them inefficient for multi-lingual environments like call centers, where accurate and rapid language detection is crucial.
Innovation Solution
A method and apparatus that combines acoustical language identification with speech-to-text processing, using estimated languages to verify text accuracy, and enhances models based on feedback, reducing manual labor and processing resources by prioritizing languages by frequency and using textual verification for confirmation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If acoustic language identification is used to identify spoken language, then the system can operate with low processing resources, but the error rate increases to around 60%
Solution Approach 1:
The patent combines acoustic language identification with textual verification to resolve the contradiction between low processing resources and high error rate. The system first uses acoustic identification to generate a candidate language list, then applies speech-to-text conversion and textual verification to validate the results, achieving both resource efficiency and accuracy.
Solution Approach 2:
The patent introduces speech-to-text conversion as an intermediary step between acoustic identification and final language determination. This intermediary process transforms acoustic data into text that can be verified against linguistic models, bridging the gap between resource-efficient acoustic analysis and accurate language identification.
2Reliability
If textual verification is used to verify language accuracy, then the error rate decreases, but processing resources including CPU, memory and time increase significantly
Solution Approach 1:
The patent applies partial verification by only performing textual verification on languages that appear in the candidate list generated by acoustic identification. This selective approach reduces the overall processing burden while maintaining verification accuracy for the most likely candidates.
Solution Approach 2:
The patent performs preliminary acoustic language identification to generate a shortlist of candidate languages before conducting textual verification. This preliminary action narrows down the scope of subsequent verification, reducing the computational resources required while maintaining high accuracy.
3Measurement precision
If speech to text engine is activated for multiple languages to verify language, then language identification accuracy improves, but the time and computational complexity increase
Solution Approach 1:
The patent segments the language identification process into distinct stages: acoustic identification to generate candidate languages, speech-to-text conversion, and textual verification. This segmentation allows the system to process only the most likely candidate languages through the resource-intensive verification stage, reducing overall processing time while maintaining accuracy.
Solution Approach 2:
The patent implements a dynamic verification process where the number of languages subjected to textual verification is adjusted based on the confidence scores from acoustic identification. High-confidence candidates undergo verification while low-confidence ones are discarded, creating a dynamic balance between accuracy and processing time.
Data Source
AI summary
In a multi-lingual environment, a method and apparatus for determining a language spoken in a speech utterance. The method and apparatus test acoustic feature vectors extracted from the utterances against acoustic models associated with one or more of the languages. Speech to text is then performed for the language indicated by the acoustic testing, followed by textual verification of the resulting text. During verification, the resulting text is processed by language specific NLP and verified against textual models associated with the language. The system is self-learning, i.e., once a language is verified or rejected, the relevant feature vectors are used for enhancing one or more acoustic models associated with one or more languages, so that acoustic determination may improve.


