Neural Network TTS Model Speech Style Characteristic Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech synthesis technologies fail to generate audio content that reflects the speaker's personality and emotions, leading to degraded effectiveness in broadcast programs, and lack intuitive user interfaces for easy generation and editing of audio content based on text styles.

Innovation Solution

A method and system that utilize an artificial neural network text-to-speech synthesis model to generate synthetic speech by determining and applying speech style characteristics, such as prosody, emotion, and context, through a user interface, allowing users to adjust settings for visual representation and adding effects like silence between sentences.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional speech synthesis technologies are used, then audio content can be generated without human recording, but the generated content does not reflect the speaker's personality and emotions

Engineering Contradiction:
Improveaudio content generation efficiencyVSAvoidspeaker personality and emotion reflection
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies parameter changes by transforming text into multiple speech style characteristics (prosody, emotion, context, situation) that are then used as input parameters for the neural network TTS model. This allows the system to generate synthetic speech that reflects speaker personality and emotions by adjusting these stylistic parameters rather than using conventional fixed-parameter synthesis

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces speech style characteristics as an intermediary layer between the input text and the TTS synthesis process. This intermediary representation captures the nuanced stylistic information needed to reflect speaker personality and emotions, bridging the gap between raw text and emotionally expressive synthetic speech

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of manufacture

If speech synthesis technology is used to produce broadcast programs, then human recording is eliminated, but the quality and authenticity of the audio content is degraded

Engineering Contradiction:
Improvebroadcast program production easeVSAvoidaudio content quality and authenticity
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The system changes the parameter representation from conventional phoneme-level or waveform-level parameters to high-level speech style characteristics (prosody, emotion, context, situation). This enables the generation of authentic-sounding broadcast content by controlling these stylistic parameters that directly influence perceived quality and authenticity

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent performs preliminary analysis to extract speech style characteristics from the input text before the actual TTS synthesis occurs. This preliminary action of determining prosody, emotion, context, and situation parameters ensures that the subsequent synthesis process can generate high-quality, authentic-sounding audio content with appropriate stylistic features

Inventive Principle:
Principle #10Preliminary action

3Extent of automation

If conventional TTS methods are used, then speech can be synthesized from text, but the user interface for generating and editing audio content is not intuitive

Engineering Contradiction:
Improvetext-to-speech synthesis automationVSAvoidaudio content generation and editing ease
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The system applies self-service by automatically analyzing the input text and determining appropriate speech style characteristics (prosody, emotion, context, situation) without requiring manual user input for each parameter. The user interface automatically generates and allows editing of audio content based on the text, making the process intuitive and easy to operate while maintaining high automation

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12183320B2Method and system for generating synthetic speech for text through user interface
Publication Date: 2024.12.31 NEOSAPIENCE INC
  • US12183320B2 patent drawing
  • US12183320B2 patent drawing
  • US12183320B2 patent drawing

AI summary

A method for generating synthetic speech for text through a user interface is provided. The method may include receiving one or more sentences, determining a speech style characteristic for the received one or more sentences, and outputting a synthetic speech for the one or more sentences that reflects the determined speech style characteristic. The one or more sentences and the determined speech style characteristic may be inputted to an artificial neural network text-to-speech synthesis model and the synthetic speech may be generated based on the speech data outputted from the artificial neural network text-to-speech synthesis model.