Personalized Voice Interaction Using User Voice Cloning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing intelligent devices provide limited personalized voice interaction experiences, often relying on pre-set responses that fail to mimic individual user interactions effectively.
Innovation Solution
An electronic device recognizes user voice information to initiate a voice conversation by imitating the voice and mannerisms of a specific user, enhancing interaction performance and providing a more realistic communication experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If existing intelligent devices use patterned voice replies based on set voice mode, then the device structure remains simple, but the interaction performance and user experience deteriorate
Solution Approach 1:
The patent applies voice cloning technology to copy the voice characteristics, tone, and mannerisms of specific users (e.g., parents, teachers) into the intelligent device. The system stores voice samples and facial expressions of target users, then generates synthetic voice responses that replicate these characteristics during conversations with children, transforming the device from using generic patterned replies to providing personalized voice copies.
Solution Approach 2:
The system changes the voice parameters dynamically based on the identified user and conversation context. By adjusting pitch, tone, speed, and facial expression parameters according to stored user profiles, the device adapts its voice output to match specific individuals' characteristics, thereby improving interaction performance without requiring complete redesign of the device architecture.
2Adaptability or versatility
If the electronic device stores voice and facial features of multiple users, then the personalization capability improves, but the data storage requirement and processing complexity increase
Solution Approach 1:
The patent segments the personalization system into distinct functional modules: voice feature extraction module, facial expression analysis module, and conversation generation module. Each module processes specific data independently (voice samples, facial images, conversation context) and outputs processed features that are combined to generate the final personalized response, reducing overall processing complexity while maintaining high personalization capability.
Solution Approach 2:
The system performs preliminary extraction and storage of voice features and facial expressions during an initial phase before actual conversations occur. User profiles are pre-built and stored in the database, allowing the device to quickly retrieve and apply relevant characteristics during real-time conversations without performing complex analysis at the moment of interaction, thus reducing processing complexity during use.
3Ease of operation
If the device imitates user voice and mannerisms in real-time, then the user engagement improves, but the computational resource consumption increases
Solution Approach 1:
The system performs voice feature extraction, facial expression analysis, and conversation pattern learning during preliminary offline processing phases. Pre-computed voice templates and facial expression models are stored and ready for quick retrieval during real-time conversations, significantly reducing the computational resources required during actual user interaction while maintaining high engagement through personalized voice imitation.
Solution Approach 2:
Instead of performing complex real-time voice synthesis from scratch, the system uses pre-generated voice copies and facial expression templates that are retrieved and adjusted minimally during conversations. This copying approach maintains user engagement through realistic voice imitation while consuming far fewer computational resources compared to generating entirely new voice responses in real-time.
Data Source
AI summary
Embodiments of this application provide a voice interaction method and an electronic device, and relate to the field of artificial intelligence AI technologies and the field of voice processing technologies. A specific solution includes: An electronic device may receive first voice information sent by a second user, and the electronic device recognizes the first voice information in response to the first voice information. The first voice information is used to request a voice conversation with a first user. The electronic device may have, on a basis that the electronic device recognizes that the first voice information is voice information of the second user, a voice conversation with the second user by imitating a voice of the first user and in a mode in which the first user has a voice conversation with the second user.


