Personalized AI Voice Recognition Model for Cross-User Communication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in understanding voice inputs from others due to differences in nationality, pronunciation, and language habits, leading to misinterpretation by voice recognition models, which fail to accurately convey the meaning of utterances in conversations across diverse user groups.

Innovation Solution

A method and device that utilize personalized artificial intelligence (AI) voice recognition models to transmit recognition information indicating the meaning of a voice input, determining abnormal situations where understanding is lacking, and providing notification messages to ensure accurate interpretation without increasing network overhead.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a standard voice recognition model is used to convert voice inputs to text, then the system can process voice data from multiple users, but the recognition accuracy deteriorates when users have different nationalities, pronunciation characteristics, and language habits

Engineering Contradiction:
Improveability to process voice data from diverse usersVSAvoidvoice recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent divides the single voice recognition model into multiple personalized models, with each model trained on data from a specific user. This segmentation allows the system to maintain versatility in handling diverse users while improving accuracy for each individual user group through specialized modeling.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by creating user-specific recognition models that are optimized for individual pronunciation characteristics, language habits, and nationalities. Each model has tailored quality suited to its specific user population, rather than using a uniform model for all users.

Inventive Principle:
Principle #3Local quality

2Reliability

If recognition information is transmitted to ensure accurate understanding, then communication quality improves, but network overhead increases

Engineering Contradiction:
Improvecommunication qualityVSAvoidnetwork data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies partial action by transmitting only the necessary recognition information when needed, rather than continuously transmitting all possible data. The system determines abnormal situations where understanding is lacking and provides notification messages only in those specific cases, avoiding unnecessary network traffic while maintaining communication quality when required.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11031000B2Method and device for transmitting and receiving audio data
Publication Date: 2021.06.08 SAMSUNG ELECTRONICS CO LTD
  • US11031000B2 patent drawing
  • US11031000B2 patent drawing
  • US11031000B2 patent drawing

AI summary

An artificial intelligence (AI) system configured to simulate functions of a human brain, such as recognition, determination, etc., by using a machine learning algorithm, such as deep learning, etc., and an application thereof. The AI system includes a method performed by a device to transmit and receive audio data to and from another device includes obtaining a voice input that is input by a first user of the device, obtaining recognition information indicating a meaning of the obtained voice input, transmitting the obtained voice input to the other device, determining whether an abnormal situation occurs, in which a second user of the other device does not understand the transmitted voice input, and transmitting the obtained recognition information to the other device, based on a result of the determination.