Speech Recognition Model Selection for Non-Native Speakers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition systems face challenges in accurately recognizing user speech when the user is not a native speaker, as they may have different intonations from native speakers, leading to reduced recognition precision.
Innovation Solution
A system that includes a processor and memory storing multiple language models for different native speakers, which performs automatic speech recognition using sample texts to select the most appropriate language model based on user speech data, improving recognition accuracy by personalizing the speech recognition model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single speech recognition model trained by native speaker intonations is used, then the system structure is simple, but the recognition accuracy for non-native speakers deteriorates
Solution Approach 1:
The patent divides the single speech recognition model into multiple language models, each trained with intonations from speakers of different native languages. The system segments the user base by native language and provides dedicated models for each group, thereby improving recognition accuracy for non-native speakers without significantly increasing overall system complexity through modular model management
Solution Approach 2:
The patent changes the training parameters of speech recognition models by using intonation data from speakers with different native languages. Each language model is trained with specific intonation characteristics corresponding to its target native language group, allowing the system to adapt recognition parameters to match diverse speaker profiles and improve accuracy
2Measurement precision
If multiple language models are stored and selected based on user speech, then recognition accuracy for non-native speakers is improved, but the device complexity increases
Solution Approach 1:
The patent performs preliminary classification of users by their native language and pre-selects appropriate language models before speech recognition occurs. By determining the user's native language group in advance and loading only the relevant language model, the system avoids the complexity of managing and switching between multiple models in real-time during speech processing
Solution Approach 2:
The system automatically detects the user's native language characteristics from speech patterns and self-selects the appropriate language model without requiring manual user configuration. This self-service mechanism reduces the operational complexity for users while maintaining high recognition accuracy by matching speakers with their optimized models
Data Source
AI summary
A system includes network interface, processor operatively connected to the network interface, and memory operatively connected to the processor. The memory stores a plurality of language models of a first language. The plurality of language models be different from one another. Each of the plurality of language models is provided for speakers whose native language is different from the first language. The memory further stores instructions that cause the processor to receive a user's speech data associated with a plurality of sample texts from an external electronic device, to perform automatic speech recognition (ASR) on the speech data using the plurality of language models, to select one language model from the plurality of language models based on results from the performed ASR, and to use the selected one language model as a default language model for speech of the user.


