Voice Synthesis Device Preset Parameter Editing for Singing Style

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

General users face difficulties in achieving a desired vocalization manner in voice synthesis as they lack knowledge on editing parameters such as pitch, velocity, and volume to replicate a specific singing or narrating style, similar to a 'retake' in human recording, without directly editing these parameters.

Innovation Solution

A voice synthesis device and method that generates sequence data based on music and lyrics information, allowing users to specify desired singing manners through preset processing content information, which includes editing parameters like velocity and volume, and outputs singing voices accordingly, with priority and evaluation mechanisms to refine the retake process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If users directly edit parameters (pitch, velocity, volume) to achieve desired vocalization manner, then the precision of controlling singing style is improved, but the ease of operation deteriorates because general users lack knowledge on how to edit parameters

Engineering Contradiction:
Improveprecision of controlling singing styleVSAvoidease of editing parameters
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent introduces an intermediary component (singing manner specification unit and processing content information) that mediates between the user's simple singing manner specification and the complex parameter editing. This intermediary automatically translates high-level singing manner descriptions into specific parameter edits, resolving the contradiction by shielding users from parameter complexity while maintaining precise control capability

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates copies of successful parameter editing patterns as preset processing content information. By storing and reusing proven parameter editing combinations associated with specific singing manners, the system allows users to achieve precise vocalization control through simple selection rather than manual parameter editing, thus improving ease of operation while maintaining precision

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If multiple pieces of sequence data are generated with different parameter edits, then the versatility of achieving various singing manners is improved, but the device complexity increases due to priority and evaluation mechanisms

Engineering Contradiction:
Improveversatility of achieving various singing mannersVSAvoidcomplexity of priority and evaluation mechanisms
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-establishing priority rules and evaluation criteria before the retake processing begins. By defining the selection logic in advance (how to compare and choose among multiple sequence data), the system achieves high versatility in handling different singing manners without requiring complex real-time decision-making mechanisms, thus reducing perceived device complexity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent utilizes parameter changes by systematically varying multiple parameters (pitch, velocity, volume, timing) to generate diverse sequence data for different singing manners. This approach achieves high versatility by exploring the parameter space, while the structured parameter editing framework keeps device complexity manageable through reusable processing templates

Inventive Principle:
Principle #35Parameter changes

3Productivity

If preset processing content information is used to automate parameter editing, then the productivity of retake process is improved, but the manufacturing precision may deteriorate due to automated editing without direct user control

Engineering Contradiction:
Improveproductivity of retake processVSAvoidprecision of parameter editing
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent implements feedback by allowing users to evaluate the generated sequence data and provide input on whether the singing manner matches their intent. This feedback loop enables the system to learn from user evaluations and refine its automated editing, maintaining high productivity while improving precision through iterative refinement based on user responses

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies dynamics by making the processing content information adaptable and modifiable based on user feedback and evaluation results. The system dynamically adjusts its automated editing behavior, allowing the precision to improve over time while maintaining high productivity through the established automated framework

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9355634B2Voice synthesis device, voice synthesis method, and recording medium having a voice synthesis program stored thereon
Publication Date: 2016.05.31 YAMAHA CORP
  • US9355634B2 patent drawing
  • US9355634B2 patent drawing
  • US9355634B2 patent drawing

AI summary

A voice synthesis device includes a sequence data generation unit configured to generate sequence data including a plurality of kinds of parameters for controlling vocalization of a voice to be synthesized based on music information and lyrics information, an output unit configured to output a singing voice based on the sequence data, and a processing content information acquisition unit configured to acquire a plurality of processing content information, associated with each of pieces of preset singing manner information. Each of the content information indicates contents of edit processing for all or part of the parameters. The sequence data generation unit generates a plurality of pieces of sequence data, and the sequence data are obtained by editing the all or part of the parameters included in the sequence data, based on the content information associated with one of the pieces of singing manner information specified by a user.