Personalized Voice Interaction Using User Voice Imitation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing intelligent devices provide limited personalized voice interaction experiences, often relying on fixed responses and failing to recognize individual users, resulting in suboptimal user engagement.
Innovation Solution
An electronic device that recognizes user voice information to initiate a voice conversation by imitating the voice and mannerisms of a specific user, enhancing interaction performance and providing a more personalized experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the electronic device uses fixed patterned voice replies based on a set voice mode, then the device complexity is reduced, but the adaptability or versatility of voice interaction deteriorates
Solution Approach 1:
The patent applies voice cloning technology to replicate the voice characteristics, tone, and mannerisms of specific users. The system creates virtual copies of user voices and uses them to generate personalized responses, enabling the device to adapt its voice output to match individual users without requiring complex real-time synthesis for each interaction scenario
Solution Approach 2:
The system dynamically adjusts voice parameters including pitch, tone, speed, and accent based on the identified user and conversation context. The voice synthesis engine adapts its output in real-time to match the characteristics of different users, transforming fixed patterned replies into dynamic, personalized voice interactions
2Ease of operation
If the electronic device provides personalized voice interaction, then the user experience is improved, but the device complexity increases
Solution Approach 1:
The system incorporates feedback mechanisms that analyze user responses and conversation flow to refine voice generation. The device monitors interaction patterns and adjusts its voice output accordingly, creating a feedback loop that improves personalization over time while managing computational resources efficiently
Solution Approach 2:
The system performs preliminary voice analysis and user identification before generating responses. By pre-processing voice characteristics and storing user profiles, the device can quickly retrieve and apply appropriate voice parameters during actual interaction, reducing real-time computational complexity
3Measurement precision
If the electronic device recognizes and imitates user voices, then the measurement precision of voice identification is improved, but the loss of information increases
Solution Approach 1:
The system extracts only the essential voice characteristics needed for identification and synthesis, such as pitch contours, tone patterns, and rhythmic features. By extracting only these key parameters rather than storing complete voice recordings, the system achieves accurate voice recognition while minimizing data retention and privacy risks
Solution Approach 2:
The patent introduces voice synthesis as an intermediary layer between voice recognition and response generation. Instead of directly analyzing and storing raw user voice data, the system converts recognized voice patterns into synthesized responses, reducing the need to store and process sensitive original voice information
Data Source
AI summary
This application provides a voice interaction method and an electronic device, and relates to the field of artificial intelligence (AI) technologies and the field of voice processing technologies. An example solution includes: An electronic device receiving first voice information sent by a second user, and the electronic device recognizing the first voice information in response to receiving the first voice information. The first voice information is used to request a voice conversation with a first user. The electronic device may have, on a basis that the electronic device recognizes that the first voice information is voice information of the second user, a voice conversation with the second user by imitating a voice of the first user and in a mode in which the first user has a voice conversation with the second user.


