Phoneme-Based Biasing for Cross-Lingual Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional automatic speech recognition (ASR) systems face challenges in recognizing context-dependent speech, particularly with foreign words and variations in accents and pronunciation, which are not effectively addressed by end-to-end models due to limited recognition candidates during beam-search decoding.
Innovation Solution
The method involves using a wordpiece-phoneme model and a biasing finite-state transducer (FST) to rescore phoneme sequences, allowing for contextual biasing of speech recognition towards terms in a biasing term list, even if those terms are from a different language than the model's training language, by mapping foreign language phonemes to the phoneme set of the ASR model, enabling recognition of foreign words without explicit training on those languages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If end-to-end models are used for speech recognition, then word error rates and latency are improved, but the ability to recognize foreign words and contextual terms deteriorates due to limited recognition candidates during beam-search decoding
Solution Approach 1:
The patent introduces phoneme sequences as an intermediary representation between acoustic features and text output. By mapping foreign language terms to phoneme sequences and using them as biasing candidates during decoding, the system enables recognition of foreign words without requiring explicit training on those languages, thus resolving the contradiction between E2E model efficiency and foreign language adaptability
Solution Approach 2:
The system performs preliminary action by pre-computing phoneme sequences for biasing terms (including foreign words) before decoding. These phoneme sequences are prepared in advance and integrated into the beam-search decoding process, allowing the model to efficiently recognize contextual terms and foreign words without compromising the speed and accuracy benefits of end-to-end processing
2Adaptability or versatility
If conventional ASR systems use independent contextual language models with n-gram WFST for biasing, then contextual recognition is improved, but device complexity and processing overhead increase
Solution Approach 1:
The patent merges the acoustic model, pronunciation model, and language model into a single end-to-end neural network. This integration eliminates the need for separate contextual LM and WFST components while maintaining contextual recognition capabilities through the phoneme-based biasing mechanism, thus reducing device complexity while preserving contextual adaptability
Solution Approach 2:
The end-to-end model performs multiple functions simultaneously: acoustic feature processing, phoneme prediction, and text generation within a single unified architecture. The phoneme-based biasing mechanism provides universal applicability across different languages and contexts without requiring language-specific models, reducing overall system complexity
3Adaptability or versatility
If phoneme sequences are used for biasing foreign language terms, then recognition of foreign words is improved, but processing time for mapping and rescoring may increase
Solution Approach 1:
Phoneme sequences for biasing terms are pre-computed and stored before decoding. This preliminary preparation eliminates the need for real-time phoneme mapping during beam-search decoding, significantly reducing processing time while maintaining the ability to recognize foreign words and contextual terms
Solution Approach 2:
The system uses phoneme sequences as a compact representation that can be copied and reused across different decoding operations. Instead of performing complex phoneme mapping during each recognition task, the pre-computed phoneme sequences are efficiently copied and applied as biasing candidates, minimizing processing overhead
Data Source
AI summary
A method includes receiving audio data encoding an utterance spoken by a native speaker of a first language, and receiving a biasing term list including one or more terms in a second language different than the first language. The method also includes processing, using a speech recognition model, acoustic features derived from the audio data to generate speech recognition scores for both wordpieces and corresponding phoneme sequences in the first language. The method also includes rescoring the speech recognition scores for the phoneme sequences based on the one or more terms in the biasing term list, and executing, using the speech recognition scores for the wordpieces and the rescored speech recognition scores for the phoneme sequences, a decoding graph to generate a transcription for the utterance.


