Personalized AI Voice Recognition Model for Cross-User Communication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in understanding voice inputs from others due to differences in nationality, pronunciation, and language habits, leading to misinterpretation by voice recognition models, which fail to accurately convey the meaning of utterances in conversations across diverse user groups.
Innovation Solution
A method and device that utilize personalized artificial intelligence (AI) voice recognition models to transmit recognition information indicating the meaning of a voice input, determining abnormal situations where understanding is lacking, and providing notification messages to ensure accurate interpretation without increasing network overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a standard voice recognition model is used to convert voice inputs to text, then the system can process voice data from multiple users, but the recognition accuracy deteriorates when users have different nationalities, pronunciation characteristics, and language habits
Solution Approach 1:
The patent divides the single voice recognition model into multiple personalized models, with each model trained on data from a specific user. This segmentation allows the system to maintain versatility in handling diverse users while improving accuracy for each individual user group through specialized modeling.
Solution Approach 2:
The patent applies local quality by creating user-specific recognition models that are optimized for individual pronunciation characteristics, language habits, and nationalities. Each model has tailored quality suited to its specific user population, rather than using a uniform model for all users.
2Reliability
If recognition information is transmitted to ensure accurate understanding, then communication quality improves, but network overhead increases
Solution Approach 1:
The patent applies partial action by transmitting only the necessary recognition information when needed, rather than continuously transmitting all possible data. The system determines abnormal situations where understanding is lacking and provides notification messages only in those specific cases, avoiding unnecessary network traffic while maintaining communication quality when required.
Data Source
AI summary
An artificial intelligence (AI) system configured to simulate functions of a human brain, such as recognition, determination, etc., by using a machine learning algorithm, such as deep learning, etc., and an application thereof. The AI system includes a method performed by a device to transmit and receive audio data to and from another device includes obtaining a voice input that is input by a first user of the device, obtaining recognition information indicating a meaning of the obtained voice input, transmitting the obtained voice input to the other device, determining whether an abnormal situation occurs, in which a second user of the other device does not understand the transmitted voice input, and transmitting the obtained recognition information to the other device, based on a result of the determination.


