Speech Recognition Model Selection by Utterance Style
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in accurately recognizing speech due to variations in utterance styles among users, as a single speech recognition model struggles to adapt to diverse speech features such as tone and intonation, leading to decreased recognition accuracy.
Innovation Solution
An AI apparatus and method that extracts utterance feature vectors from user speech data, determines the corresponding utterance style, and generates a speech recognition model tailored to that style, allowing for the learning of new models when significant deviations are detected, thereby improving recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single speech recognition model is used, then device complexity is reduced, but speech recognition accuracy deteriorates due to inability to adapt to various speech styles
Solution Approach 1:
The patent divides a single speech recognition model into multiple style-specific models (first speech recognition model for formal style, second speech recognition model for informal style). Each model is specialized for recognizing speech in a particular utterance style, thereby improving recognition accuracy for different speech patterns without requiring a single complex universal model.
Solution Approach 2:
The system changes the parameter of model selection based on detected utterance style. When formal style is detected, the first speech recognition model is used; when informal style is detected, the second speech recognition model is used. This dynamic parameter adjustment optimizes recognition accuracy according to the specific speech style being analyzed.
2Measurement precision
If multiple speech recognition models are used for different utterance styles, then speech recognition accuracy is improved, but device complexity increases
Solution Approach 1:
The patent introduces an utterance style detection model as an intermediary between the speech input and the multiple speech recognition models. This mediator detects the utterance style first and then directs the speech signal to the appropriate recognition model, managing the complexity of having multiple models by providing a systematic selection mechanism.
Solution Approach 2:
The system dynamically selects which speech recognition model to use based on the detected utterance style. The model selection is not fixed but adapts in real-time according to the speech style characteristics, allowing the system to optimize performance for each specific input while managing complexity through adaptive rather than static model configuration.
3Measurement precision
If a speech recognition model is trained only on specific utterance styles, then recognition accuracy for those styles is improved, but adaptability to new utterance styles deteriorates
Solution Approach 1:
The system incorporates feedback through the utterance style detection model that continuously monitors incoming speech and identifies the utterance style. This feedback mechanism allows the system to adapt to different speech styles by detecting their characteristics and selecting or training appropriate recognition models, thereby maintaining both accuracy for known styles and adaptability to new styles.
Data Source
AI summary
Disclosed herein an artificial intelligence apparatus for recognizing speech in consideration of an utterance style including a microphone, and a processor configured to obtain, via the microphone, speech data including speech of a user, extract an utterance feature vector from the obtained speech data, determine an utterance style corresponding to the speech based on the extracted utterance feature vector, and generate a speech recognition result using a speech recognition model corresponding to the determined utterance style.


