Automated TTS Voice Development via Crowdsourced Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The development of languages and voices for text-to-speech (TTS) systems is time-consuming and requires specialized knowledge, as existing methods rely heavily on human evaluation and linguistic expertise, limiting efficiency and scalability.

Innovation Solution

A system that automates the development of languages and voices for TTS systems using machine learning algorithms and crowdsourced feedback, allowing for recursive testing and modification of conversion rules and speech segments, reducing the need for human intervention and expertise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If human evaluation and linguistic expertise are used to develop languages and voices for TTS systems, then the accuracy and quality of synthesized speech is improved, but the development time and resource requirements increase significantly

Engineering Contradiction:
Improvequality of synthesized speechVSAvoiddevelopment time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system implements automated feedback loops where synthetic speech is evaluated by listeners and the results are used to iteratively improve conversion rules and speech segments. This automated feedback mechanism replaces manual human evaluation, maintaining quality improvement while dramatically reducing development time and resource requirements.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The TTS development system performs self-evaluation and self-improvement through automated processes. The system autonomously generates test sentences, evaluates synthetic speech quality, identifies errors in conversion rules, and modifies speech segments without requiring continuous human intervention, enabling rapid iterative development.

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If specialized linguistic expertise is required for developing TTS languages and voices, then the accuracy of language conversion rules is improved, but the complexity and difficulty of the development process increases

Engineering Contradiction:
Improveaccuracy of conversion rulesVSAvoiddevelopment process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system replaces the mechanical process of manual linguistic analysis and rule creation with automated computational processes. Machine learning algorithms automatically analyze speech data, identify conversion patterns, and generate language rules, eliminating the need for specialized linguistic expertise while maintaining rule accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces automated evaluation metrics and intermediate processing layers that bridge the gap between raw speech data and final conversion rules. These intermediaries automatically handle the complex analysis and rule generation tasks, simplifying the overall development process while preserving accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If manual testing and evaluation of TTS systems is performed, then the quality control and error detection are improved, but the productivity and scalability of development are reduced

Engineering Contradiction:
Improvequality controlVSAvoiddevelopment efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system implements continuous automated testing and evaluation that operates without interruption. Test sentences are continuously generated, synthetic speech is continuously evaluated, and conversion rules are continuously refined, maintaining high quality control while enabling rapid iterative development and improving overall productivity.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system performs preliminary automated evaluation and error detection before manual review is needed. By pre-identifying and correcting obvious errors in conversion rules and speech segments through automated processes, the system maintains quality control while reducing the time and resources required for manual testing and evaluation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9196240B2Automated text to speech voice development
Publication Date: 2015.11.24 AMAZON TECH INC
  • US9196240B2 patent drawing
  • US9196240B2 patent drawing
  • US9196240B2 patent drawing

AI summary

A group of users may be presented with text and a synthesized speech recording of the text. The users can listen to the synthesized speech recording and submit feedback regarding errors or other issues with the synthesized speech. A system of one or more computing devices can analyze the feedback, modify the voice or language rules, and recursively test the modifications. The modifications may be determined through the use of machine learning algorithms or other automated processes.