Voice Synthesizer Real-Time Vocalization Timing Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice synthesis techniques cannot vary vocalization time points in real-time, limiting user control over the timing of phoneme vocalization in generated voices.

Innovation Solution

A voice synthesizing method and apparatus that determines a manipulation position based on user input, allowing vocalization of a first phoneme to start before reaching a reference position and complete by the time the manipulation position reaches that position, enabling real-time control over the transition to a second phoneme.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If preset vocalization time points are used for respective notes, then voice synthesis can be performed with fixed timing, but the vocalization time points cannot be varied on a real-time basis

Engineering Contradiction:
Improvereal-time control of vocalization timingVSAvoidcontrol mechanism complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic control of vocalization time points by allowing the manipulation position to move along the time axis in real-time according to user input. The voice synthesis unit adjusts the timing of phoneme vocalization dynamically during voice generation, transforming fixed preset time points into variable real-time controllable parameters. This enables adaptive control without requiring complex preprocessing or multiple static configurations.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If vocalization time points are fixed in advance, then voice generation is simple and fast, but user control over timing is limited

Engineering Contradiction:
Improveuser control capabilityVSAvoidprocessing delay
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent performs preliminary determination of the manipulation position based on user input before actual voice generation. The system calculates the relationship between the manipulation position and reference position in advance, and uses this pre-computed information to control vocalization timing during voice generation. This preliminary action enables real-time control without introducing significant processing delays during the actual voice synthesis process.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the manipulation position is determined based on user manipulation, then real-time variation of vocalization time points is enabled, but the system requires continuous monitoring and adjustment

Engineering Contradiction:
Improvedynamic timing adjustmentVSAvoidmonitoring and adjustment mechanism
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism where the voice synthesis unit continuously monitors the manipulation position and adjusts the vocalization time points accordingly. The system calculates the difference between the manipulation position and reference position, and uses this feedback information to dynamically control the timing of phoneme transitions. This feedback loop enables automatic real-time adjustment without requiring complex external monitoring systems.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9424831B2Voice synthesizing having vocalization according to user manipulation
Publication Date: 2016.08.23 YAMAHA CORP
  • US9424831B2 patent drawing
  • US9424831B2 patent drawing
  • US9424831B2 patent drawing

AI summary

A voice synthesizing apparatus includes a manipulation determiner configured to determine a manipulation position which is moved according to a manipulation of a user, and a voice synthesizer configured to generate, in response to an instruction to generate a voice in which a second phoneme follows a first phoneme, a voice signal so that vocalization of the first phoneme starts before the manipulation position reaches a reference position and that vocalization from the first phoneme to the second phoneme is made when the manipulation position reaches the reference position.