Language Adaptivity in Speech Recognition Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multilingual speech recognition systems face low accuracy due to confusion between pronunciations in different languages, leading to a worse user experience and reduced processing efficiency.
Innovation Solution
A method and apparatus for speech recognition based on language adaptivity, which involves extracting phoneme features from voice data, using a pre-trained language discrimination model to determine the language, and switching to the corresponding language acoustic model for accurate recognition, thereby avoiding confusion between languages.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a hybrid acoustic pronunciation unit set including multiple languages is used to train an acoustic model, then the speech recognition system can support multiple languages, but the recognition accuracy in different languages is greatly affected due to pronunciation confusion
Solution Approach 1:
The patent segments the multilingual speech recognition system into separate language-specific acoustic models and a language discrimination model. Instead of using a single hybrid acoustic model that mixes pronunciations from multiple languages, the system divides the recognition process into two stages: first identifying the language through the discrimination model, then routing to the appropriate language-specific acoustic model. This segmentation eliminates pronunciation confusion while maintaining multilingual support.
Solution Approach 2:
The patent introduces a language discrimination model as an intermediary component between the voice input and the acoustic models. This intermediary first analyzes the phoneme features of the input voice data to determine the language type, then directs the recognition process to the corresponding language-specific acoustic model. This intermediary structure prevents direct mixing of different language pronunciations while enabling seamless multilingual recognition.
2Measurement precision
If manual language selection is required for speech recognition, then accurate language-specific recognition can be achieved, but unnecessary user operations reduce processing efficiency
Solution Approach 1:
The patent enables the speech recognition system to automatically identify and select the appropriate language without user intervention. The language discrimination model autonomously analyzes the input voice data's phoneme features to determine the language type, then automatically routes to the corresponding acoustic model. This self-service mechanism eliminates the need for manual language selection while maintaining accurate language-specific recognition, thereby improving processing efficiency.
Solution Approach 2:
The patent performs language identification as a preliminary action before executing the main speech recognition task. By first analyzing the phoneme features and determining the language type through the discrimination model, the system prepares the appropriate language-specific acoustic model in advance. This preliminary language identification action enables subsequent accurate recognition without requiring user input, streamlining the overall processing workflow.
Data Source
AI summary
A method for speech recognition based on language adaptivity comprises obtaining voice data of a user. The method also comprises extracting, based on the obtained voice data, a phoneme feature representing pronunciation phoneme information. The phoneme feature is input to a pre-trained language discrimination model that is pre-trained based on a multilingual corpus. A language discrimination result corresponding to the phoneme feature and in accordance with the language discrimination model is obtained. The method also comprises obtaining a speech recognition result of the voice data based on a language acoustic model of a language corresponding to the language discrimination result. The method further comprises determining a speech recognition result of the voice data based on a language acoustic model of a language corresponding to the language discrimination result.


