Speech Synthesis Model Training With Targeted Text Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-speech synthesis systems require extensive human training data, often provided by professional actors, to achieve realistic and expressive speech, leading to inefficiencies in training time and resource utilization.

Innovation Solution

A computer-implemented method for training and testing a speech synthesis model that automatically identifies performance gaps and requests targeted training data from actors, reducing the need for extensive training by focusing on specific areas where the model is underperforming.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If professional actors provide extensive training data to achieve realistic and expressive speech output, then the quality of speech synthesis is improved, but the training time and resource requirements increase significantly

Engineering Contradiction:
Improvespeech qualityVSAvoidtraining time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system automatically evaluates its own performance on test sentences and identifies specific areas where improvement is needed, then autonomously generates targeted training data requests without requiring continuous human supervision or assessment

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes the parameters of training data collection by shifting from extensive random data collection to targeted data collection based on specific performance metrics and identified weaknesses, thereby improving efficiency while maintaining quality

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If professional actors are used to provide training data for realistic speech, then the expressiveness of the output is improved, but the cost and complexity of the training process increase

Engineering Contradiction:
Improvespeech expressivenessVSAvoidtraining process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system implements a feedback loop where test results automatically inform training data requirements, creating a closed-loop system that reduces complexity by eliminating manual assessment steps while maintaining high expressiveness through targeted training

Inventive Principle:
Principle #23Feedback

3Manufacturing precision

If extensive training data is collected to improve speech synthesis quality, then the realism of output speech is improved, but the quantity of required training data increases

Engineering Contradiction:
Improvespeech realismVSAvoidtraining data volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The system applies local quality by focusing training data collection on specific sentences and performance areas identified as needing improvement, rather than uniformly collecting extensive data across all possible speech scenarios, thereby achieving realism with reduced overall data volume

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12592217B2System and method for speech processing
Publication Date: 2026.03.31 SPOTIFY
  • US12592217B2 patent drawing
  • US12592217B2 patent drawing
  • US12592217B2 patent drawing

AI summary

A method for training a speech synthesis model adapted to output speech in response to input text is provided. The method includes receiving training data for training said speech synthesis model, the training data comprising speech that corresponds to known text. The method includes training said speech synthesis model. The method includes testing said speech synthesis model using a plurality of text sequences. The method includes calculating at least one metric indicating the performance of the model when synthesising each text sequence. The method includes determining from said metric whether the speech synthesis model requires further training. The method includes determining targeted training text from said calculated metrics, wherein said targeting training text is text related to text sequences where the metric indicated that the model required further training. And the method includes outputting said determined targeted training text with a request further speech corresponding to the targeted training text.