Personal Assistant Voice Adaptation via Pre-Extracted Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI personal assistant systems lack the ability to dynamically adjust their voice output to better match user preferences, leading to suboptimal user experience in terms of voice interaction.

Innovation Solution

The system and method involve an electronic device that extracts voice data features from media content, transmits these features to a server, which generates and transmits voice data to the device to output a personalized voice, allowing the AI personal assistant to adapt its voice to user preferences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the AI personal assistant uses a fixed voice output, then the system complexity is low, but the adaptability to user preferences is poor

Engineering Contradiction:
Improveadaptability to user preferencesVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system pre-extracts voice data features from media content and stores them in advance. When a user requests a voice change, the pre-processed features are quickly retrieved and applied, rather than processing from scratch. This preliminary preparation reduces the complexity of real-time voice adaptation while maintaining high adaptability to user preferences.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces voice data features as an intermediary between the media content and the AI personal assistant's voice output. These features serve as a bridge that enables the assistant to adapt its voice without requiring complex direct analysis of media content during interaction, thus reducing system complexity while improving adaptability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the system processes voice data in real-time during interaction, then the adaptability is high, but the response time increases

Engineering Contradiction:
Improvevoice adaptation capabilityVSAvoidresponse time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

Voice data features are extracted and prepared in advance during media playback, before the user needs to interact with the personal assistant. This preliminary processing ensures that when voice adaptation is requested, the system can immediately apply pre-computed features without delaying the user interaction, thus maintaining both high adaptability and fast response time.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If the AI personal assistant provides generic voice output, then the ease of operation is high, but the user experience quality is low

Engineering Contradiction:
Improveease of voice interactionVSAvoiduser experience quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system applies different voice characteristics to different parts of the interaction based on user preferences and context. Instead of using a uniform generic voice, the AI personal assistant adapts specific voice attributes (such as pitch, tone, or timbre) to match user preferences, thereby maintaining ease of operation while significantly improving user experience quality through localized voice customization.

Inventive Principle:
Principle #3Local quality

Data Source

PatentEP3762819B1Electronic device and method of controlling thereof
Publication Date: 2023.08.02 SAMSUNG ELECTRONICS CO LTD
  • EP3762819B1 patent drawingFigure 1~2
  • EP3762819B1 patent drawingFigure 3~4
  • EP3762819B1 patent drawingFigure 5

AI summary

An electronic device for changing a voice of a personal assistant function, and a method therefor are provided. The electronic device includes a display, a transceiver, processor, and a memory for storing commands executable by the processor. The processor is configured to, based on a user command to request acquisition of voice data feature of a person included in a media content displayed on the display being received, control the display to display information of a person, based on a user input to select the one of the information of a person being received, acquire voice data corresponding to an utterance of a person related to the selected information of a person, and acquire voice data feature from the acquired voice data, control the transceiver to transmit the acquired voice data feature to a server.