Personalized Text-to-Speech Training With Intermediate User Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Generating personalized text-to-speech models requires significant computational resources and time, and the performance of such models is often not meeting user expectations due to the lack of effective feedback mechanisms during the training process.

Innovation Solution

An electronic device incorporates a feedback mechanism that allows users to provide input on intermediate results during the training process, adjusting the model to improve performance and reduce generation time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If deep learning algorithms are used to generate P-TTS models, then the model can achieve personalized speech generation, but the computational resources and time required increase significantly

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the training process into multiple stages: initial model training with base data, intermediate evaluation with user feedback, and final refinement. This allows the system to provide usable intermediate results during training rather than requiring complete training before any useful output, thereby reducing the perceived time loss while maintaining model performance.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a feedback mechanism where users can evaluate intermediate training results and provide feedback that guides further training. This feedback loop allows the system to optimize training efficiency by focusing computational resources on areas that most improve user satisfaction, reducing overall training time while maintaining high model performance.

Inventive Principle:
Principle #23Feedback

2Reliability

If deep learning algorithms are used to generate P-TTS models, then the model can achieve personalized speech generation, but the computational resources required increase significantly

Engineering Contradiction:
Improvemodel performanceVSAvoidcomputational resources
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by training an intermediate model with a portion of the training data to provide initial results, rather than requiring complete training with all data. This allows the system to deliver functional P-TTS capabilities with reduced computational resource consumption, while offering the option to continue training for improved performance.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system dynamically adjusts training parameters such as learning rate, batch size, and training epochs based on user feedback and intermediate results. This allows optimization of computational resource usage by stopping training when sufficient performance is achieved, rather than consuming fixed high computational resources regardless of actual need.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If the training process is extended to improve model performance, then better P-TTS models are generated, but the time consumption increases

Engineering Contradiction:
Improvemodel performanceVSAvoidgeneration efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary training to generate an intermediate model that can already produce functional P-TTS results. This preliminary action provides immediate value to users while the system continues background training to improve performance, thereby maintaining high productivity without sacrificing eventual model quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The training process continues in the background while intermediate results are already being used. This continuity ensures that the system always provides useful P-TTS generation capability while progressively improving model performance, maintaining high productivity throughout the entire training lifecycle.

Inventive Principle:
Principle #20Continuity of useful action

4Reliability

If user feedback is incorporated during training, then model performance improves, but the complexity of the training process increases

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a structured feedback mechanism where users evaluate intermediate results through simple interfaces. The feedback is automatically processed to adjust training parameters, improving model performance without requiring complex manual intervention. The feedback system itself is designed to be simple and intuitive, minimizing the complexity burden on users.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system automatically processes user feedback and adjusts training parameters without requiring complex user configuration or intervention. The feedback loop is self-managing, where the system interprets user responses and autonomously modifies the training process, thereby improving performance while keeping the user-facing complexity low.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP4310835B1Electronic device and personalized text-to-speech model generation method by electronic device
Publication Date: 2026.03.04 SAMSUNG ELECTRONICS CO LTD
  • EP4310835B1 patent drawingFigure 1
  • EP4310835B1 patent drawingFigure 2
  • EP4310835B1 patent drawingFigure 3

AI summary

An electronic device includes a memory storing instructions and a processor configured to execute the instructions. When the instructions are executed by the processor, the processor records a speech of a user corresponding to a text and obtains recorded data in which the text and the speech of the user are matched, stores an intermediate model trained based on a portion of the recorded data while training a speech model to generate a personalized text-to-speech (P-TTS) model corresponding to the user, generates an intermediate result from the training using the intermediate model and provides the generated intermediate result to the user, and receives feedback from the user on the intermediate result. Other example embodiments, in addition to the foregoing example embodiment, are also applicable.