Multi-language Speech Recognition System with Mixed Dictionary
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multi-language mixed speech recognition systems face challenges in accurately segmenting mixed speech and maintaining recognition accuracy due to short context segments and the requirement for unusual pronunciation of proper nouns, leading to reduced overall sentence recognition accuracy.
Innovation Solution
A multi-language mixed speech recognition method is developed, involving the configuration of a mixed dictionary with language marks, training with multi-language speech and text data to form acoustic and language recognition models, and adjusting the output layer probabilities to improve recognition accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If individual speech recognition systems are established for each language and mixed speech is segmented, then the system can recognize different languages separately, but the segmentation accuracy is difficult to ensure and context information becomes too short affecting recognition accuracy
Solution Approach 1:
The patent merges multiple individual speech recognition systems into a single unified multi-language speech recognition system. Instead of treating each language separately with individual recognition systems, the invention combines them into one system that can handle multiple languages simultaneously, thereby preserving context information across language boundaries and improving segmentation accuracy through unified processing.
Solution Approach 2:
The patent creates a universal speech recognition system that can recognize multiple languages through a single interface. The system uses a unified acoustic model and language model that adapt to different languages, eliminating the need for separate recognition systems for each language while maintaining the ability to distinguish and process different languages appropriately.
2Adaptability or versatility
If dictionary expansion is performed by piecing together another language using phone sets of one language, then vocabularies of different languages can be recognized, but the user must pronounce in a very strange manner and the accuracy of recognizing the entire sentence is greatly reduced
Solution Approach 1:
The patent changes the fundamental parameters of the speech recognition system by using language-specific phone sets and pronunciation models for each language instead of forcing all languages to use the same phone set. The system adjusts acoustic models and language models according to the specific characteristics of each language, allowing users to pronounce words naturally in their own language while maintaining high recognition accuracy.
Solution Approach 2:
The patent applies segmentation at the language level rather than forcing phonetic transcription across languages. The system segments speech by language boundaries and processes each language segment with its own optimized models, preserving the natural pronunciation of each language while maintaining overall sentence recognition accuracy.
3Measurement precision
If a single unified speech recognition system is used for multiple languages, then context information is preserved, but the system must handle the complexity of distinguishing and processing multiple languages simultaneously
Solution Approach 1:
The patent introduces language models and acoustic models as intermediary components that mediate between the input speech and the recognition output. These intermediary models are trained to recognize language-specific patterns and characteristics, automatically distinguishing between different languages and routing processing appropriately, thereby managing the complexity of multi-language processing while preserving context information.
Data Source
AI summary
The invention discloses a multi-language mixed speech recognition method, which belongs to the technical field of speech recognition; the method comprises: step S1, configuring a multi-language mixed dictionary including a plurality of different languages; step S2, performing training according to the multi-language mixed dictionary and multi-language speech data including a plurality of different languages to form an acoustic recognition model; step S3, performing training according to multi-language text corpus including a plurality of different languages to form a language recognition model; step S4, forming the speech recognition system by using the multi-language mixed dictionary, the acoustic recognition model and the language recognition model; and subsequently, recognizing mixed speech by using the speech recognition system, and outputting a corresponding recognition result. The above technical solution has the beneficial effects of being able to support the recognition of mixed speech in multiple languages, improving the accuracy and efficiency of recognition, and thus improving the performance of the speech recognition system.


