Personalized Voice Interaction Using User Voice Cloning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing intelligent devices provide limited personalized voice interaction experiences, often relying on pre-set responses that fail to mimic individual user interactions effectively.

Innovation Solution

An electronic device recognizes user voice information to initiate a voice conversation by imitating the voice and mannerisms of a specific user, enhancing interaction performance and providing a more realistic communication experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If existing intelligent devices use patterned voice replies based on set voice mode, then the device structure remains simple, but the interaction performance and user experience deteriorate

Engineering Contradiction:
Improveinteraction performanceVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent applies voice cloning technology to copy the voice characteristics, tone, and mannerisms of specific users (e.g., parents, teachers) into the intelligent device. The system stores voice samples and facial expressions of target users, then generates synthetic voice responses that replicate these characteristics during conversations with children, transforming the device from using generic patterned replies to providing personalized voice copies.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the voice parameters dynamically based on the identified user and conversation context. By adjusting pitch, tone, speed, and facial expression parameters according to stored user profiles, the device adapts its voice output to match specific individuals' characteristics, thereby improving interaction performance without requiring complete redesign of the device architecture.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If the electronic device stores voice and facial features of multiple users, then the personalization capability improves, but the data storage requirement and processing complexity increase

Engineering Contradiction:
Improvepersonalization capabilityVSAvoiddata processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the personalization system into distinct functional modules: voice feature extraction module, facial expression analysis module, and conversation generation module. Each module processes specific data independently (voice samples, facial images, conversation context) and outputs processed features that are combined to generate the final personalized response, reducing overall processing complexity while maintaining high personalization capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary extraction and storage of voice features and facial expressions during an initial phase before actual conversations occur. User profiles are pre-built and stored in the database, allowing the device to quickly retrieve and apply relevant characteristics during real-time conversations without performing complex analysis at the moment of interaction, thus reducing processing complexity during use.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If the device imitates user voice and mannerisms in real-time, then the user engagement improves, but the computational resource consumption increases

Engineering Contradiction:
Improveuser engagementVSAvoidcomputational resource consumption
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The system performs voice feature extraction, facial expression analysis, and conversation pattern learning during preliminary offline processing phases. Pre-computed voice templates and facial expression models are stored and ready for quick retrieval during real-time conversations, significantly reducing the computational resources required during actual user interaction while maintaining high engagement through personalized voice imitation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of performing complex real-time voice synthesis from scratch, the system uses pre-generated voice copies and facial expression templates that are retrieved and adjusted minimally during conversations. This copying approach maintains user engagement through realistic voice imitation while consuming far fewer computational resources compared to generating entirely new voice responses in real-time.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20260112367A1Voice interaction method and electronic device
Publication Date: 2026.04.23 HUAWEI TECH CO LTD
  • US20260112367A1 patent drawing
  • US20260112367A1 patent drawing
  • US20260112367A1 patent drawing

AI summary

Embodiments of this application provide a voice interaction method and an electronic device, and relate to the field of artificial intelligence AI technologies and the field of voice processing technologies. A specific solution includes: An electronic device may receive first voice information sent by a second user, and the electronic device recognizes the first voice information in response to the first voice information. The first voice information is used to request a voice conversation with a first user. The electronic device may have, on a basis that the electronic device recognizes that the first voice information is voice information of the second user, a voice conversation with the second user by imitating a voice of the first user and in a mode in which the first user has a voice conversation with the second user.