Personalized Speech Recognition via User Image Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing artificial intelligence personal assistant systems lack user-specific responses due to generic training of neural network models, leading to undesired or unsuitable responses to user speech.
Innovation Solution
An electronic apparatus and method that obtain user information through image recognition and select a trained neural network model corresponding to the user, allowing for personalized speech recognition, natural language understanding, and text-to-speech responses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a plurality of neural network models are generically trained without considering user characteristics, then the system complexity is reduced and ease of manufacture is improved, but the adaptability to different users deteriorates and usability is limited
Solution Approach 1:
The patent segments the generic neural network model into multiple user-specific models by introducing a user identification module that divides users into different groups (e.g., children, adults, seniors) and selects appropriate models for each group. This allows the system to maintain multiple specialized models without requiring complex individual training for each user.
Solution Approach 2:
The patent applies preliminary action by pre-training multiple neural network models for different user characteristics before actual use. The system prepares various speech recognition models, natural language understanding models, and text-to-speech models in advance for different user groups, so that when a user interacts with the system, the appropriate pre-trained model is already available for immediate use.
2Adaptability or versatility
If user-specific neural network models are implemented to provide personalized responses, then the adaptability to user characteristics is improved, but the device complexity increases
Solution Approach 1:
The patent introduces an intermediary module (user identification module and model selection module) that mediates between the user input and the multiple neural network models. This intermediary layer manages the complexity by automatically identifying user characteristics and selecting the appropriate pre-trained models, shielding the user from the underlying system complexity.
Solution Approach 2:
The patent creates universal model families that can serve multiple user groups. Instead of creating entirely separate systems for each user, the system uses universal architectures (speech recognition, natural language understanding, text-to-speech) that are trained with different parameters for different user groups, reducing overall system complexity while maintaining adaptability.
3Ease of operation
If generic neural network models are used for speech recognition and response generation, then the ease of operation is maintained, but the response accuracy and user satisfaction deteriorate
Solution Approach 1:
The patent applies local quality by tailoring the neural network model characteristics to match specific user groups' speech patterns and preferences. For example, children's speech recognition models are trained on child-specific vocabulary and speech patterns, while adult models are optimized for adult speech characteristics, thereby improving recognition accuracy for each local user group without complicating the overall system operation.
Data Source
AI summary
An electronic apparatus and a control method thereof are provided. The electronic apparatus includes a microphone, a camera, a memory storing an instruction, and a processor configured to control the electronic apparatus coupled with the microphone, the camera and the memory, and the processor is configured to, by executing the instruction, obtain a user image by photographing a user through the camera, obtain the user information based on the user image, and based on a user speech being input from the user through the microphone, recognize the user speech by using a speech recognition model corresponding to the user information among a plurality of speech recognition models.


