Regional Dialect Phoneme Adaptive Training System
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems are inadequate in recognizing regional dialects as they are often based on standard dialects, leading to reduced recognition performance and requiring manual transcription, which is time-consuming and costly.
Innovation Solution
A regional dialect phoneme adaptive training system and method that processes regional dialect speech by transcribing text data, generating a regional dialect corpus, and training phoneme adaptive models using extracted phonemes and frequencies, allowing for improved recognition without converting regional dialects to standard dialects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a speech recognition system is created based on the standard dialect, then the recognition accuracy for standard dialect is improved, but the recognition capability for regional dialect is significantly reduced
Solution Approach 1:
The patent changes the parameter of dialect type from standard dialect to regional dialect by extracting and training on regional dialect phonemes. The system extracts phonemes from regional dialect speech data and uses these extracted phonemes to train the acoustic model, thereby adapting the model to recognize regional dialect characteristics while maintaining standard dialect recognition capability.
Solution Approach 2:
The patent segments the speech recognition task into separate phoneme extraction and model training stages. It extracts phonemes from regional dialect speech data separately, then uses these extracted phonemes to train the acoustic model. This segmentation allows the system to handle regional dialect variations without compromising standard dialect recognition.
2Measurement precision
If manual transcription is performed for speech data processing, then the accuracy of text data is improved, but the time consumption and cost increase significantly
Solution Approach 1:
The patent implements self-service through automatic phoneme extraction and text data generation. The system automatically extracts phonemes from speech data and generates text data without requiring manual transcription by humans. This automation maintains high accuracy while significantly reducing time consumption and operational costs.
Solution Approach 2:
The patent replaces the mechanical manual transcription process with an automated computational system. Instead of human transcribers manually converting speech to text, the system uses automated phoneme extraction algorithms and text generation processes, thereby eliminating manual labor while maintaining accuracy.
3Device complexity
If regional dialect speech is converted to standard dialect speech for recognition, then the recognition process is simplified, but the ability to distinguish and accurately recognize regional dialect characteristics is lost
Solution Approach 1:
Instead of converting regional dialect to standard dialect for recognition, the patent inverts the approach by extracting regional dialect phonemes and training the model specifically on these phonemes. This allows the system to recognize regional dialect characteristics directly without conversion, maintaining both simplicity and accuracy.
Solution Approach 2:
The patent changes the parameter of the acoustic model from standard dialect-based to regional dialect-based by using extracted regional dialect phonemes for training. This parameter change enables the model to accurately distinguish and recognize regional dialect characteristics while keeping the recognition process simple.
Data Source
AI summary
Disclosed are a regional dialect phoneme adaptive training method and system. The regional dialect phoneme adaptive training method includes transcription of text data, and generation of a regional dialect corpus based on the text data and regional dialect-containing speech data, and generation of an acoustic model and a language model using the regional dialect corpus. The generation of an acoustic model and a language model may be performed by machine learning of an artificial intelligence (AI) algorithm in which phonemes of a regional dialect item and a frequency of the phonemes of the regional dialect item are extracted and used. A user is able to use a regional dialect speech recognition service which is improved using 5G mobile communication technologies of eMBB, URLLC, or mMTC.


