Voice Synthesis Apparatus Using Multi-Model Control Data Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice synthesis techniques struggle to accurately reflect a user's intention or taste in the synthesized voice, as they often lack the ability to effectively adjust phonetic identifiers, pitch, and sound periods, leading to a lack of personalization in the generated voice.

Innovation Solution

A voice synthesis method and apparatus that utilize multiple trained models to generate control data representing frequency characteristics of a voice, allowing users to input instructions to adjust phonetic identifiers, pitch, and sound periods, resulting in a voice signal that better suits their intended style and taste.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional phoneme-based voice synthesis techniques are used, then voice generation is achieved, but the ability to accurately reflect user intention or taste is insufficient

Engineering Contradiction:
Improveaccuracy of reflecting user intentionVSAvoiduser customization capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the voice synthesis process into multiple independent controllable dimensions: phoneme identification, pitch control, sound period adjustment, and timbre modification. Each dimension can be independently adjusted according to user instructions, allowing precise control over different aspects of the synthesized voice to accurately reflect user intention.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts multiple voice parameters (pitch, sound period, timbre) based on real-time user instructions. The synthesis process is not static but adapts continuously to user feedback, allowing the generated voice to evolve and match user taste through iterative refinement of control data.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If multiple trained models are used to generate control data, then voice synthesis accuracy improves, but system complexity increases

Engineering Contradiction:
Improvevoice synthesis accuracyVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple specialized trained models (phoneme identification model, pitch control model, sound period model) into an integrated voice synthesis system. These models work together in a coordinated manner, where each model handles a specific aspect of voice generation, and their outputs are merged to produce the final synthesized voice with high accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system employs universal trained models that can handle multiple functions within the voice synthesis process. For example, the phoneme identification model not only identifies phonemes but also provides basis for pitch and timing control. This multi-functionality reduces the need for completely separate specialized models for each task.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If user instructions are incorporated to adjust control data, then personalization of synthesized voice improves, but processing time increases

Engineering Contradiction:
Improvevoice personalizationVSAvoidsynthesis processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary processing of user instructions by pre-training multiple specialized models offline. These models are prepared in advance with pre-learned parameters and characteristics, so that during actual voice synthesis, the system can quickly apply pre-computed control data adjustments without extensive real-time computation, reducing processing time while maintaining personalization.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11495206B2Voice synthesis method, voice synthesis apparatus, and recording medium
Publication Date: 2022.11.08 YAMAHA CORP
  • US11495206B2 patent drawing
  • US11495206B2 patent drawing
  • US11495206B2 patent drawing

AI summary

Voice synthesis method and apparatus generate second control data using an intermediate trained model with first input data including first control data designating phonetic identifiers, change the second control data in accordance with a first user instruction provided by a user, generate synthesis data representing frequency characteristics of a voice to be synthesized using a final trained model with final input data including the first control data and the changed second control data, and generate a voice signal based on the generated synthesis data.