Voice Synthesis Device Preset Parameter Editing for Singing Style
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
General users face difficulties in achieving a desired vocalization manner in voice synthesis as they lack knowledge on editing parameters such as pitch, velocity, and volume to replicate a specific singing or narrating style, similar to a 'retake' in human recording, without directly editing these parameters.
Innovation Solution
A voice synthesis device and method that generates sequence data based on music and lyrics information, allowing users to specify desired singing manners through preset processing content information, which includes editing parameters like velocity and volume, and outputs singing voices accordingly, with priority and evaluation mechanisms to refine the retake process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If users directly edit parameters (pitch, velocity, volume) to achieve desired vocalization manner, then the precision of controlling singing style is improved, but the ease of operation deteriorates because general users lack knowledge on how to edit parameters
Solution Approach 1:
The patent introduces an intermediary component (singing manner specification unit and processing content information) that mediates between the user's simple singing manner specification and the complex parameter editing. This intermediary automatically translates high-level singing manner descriptions into specific parameter edits, resolving the contradiction by shielding users from parameter complexity while maintaining precise control capability
Solution Approach 2:
The patent creates copies of successful parameter editing patterns as preset processing content information. By storing and reusing proven parameter editing combinations associated with specific singing manners, the system allows users to achieve precise vocalization control through simple selection rather than manual parameter editing, thus improving ease of operation while maintaining precision
2Adaptability or versatility
If multiple pieces of sequence data are generated with different parameter edits, then the versatility of achieving various singing manners is improved, but the device complexity increases due to priority and evaluation mechanisms
Solution Approach 1:
The patent applies preliminary action by pre-establishing priority rules and evaluation criteria before the retake processing begins. By defining the selection logic in advance (how to compare and choose among multiple sequence data), the system achieves high versatility in handling different singing manners without requiring complex real-time decision-making mechanisms, thus reducing perceived device complexity
Solution Approach 2:
The patent utilizes parameter changes by systematically varying multiple parameters (pitch, velocity, volume, timing) to generate diverse sequence data for different singing manners. This approach achieves high versatility by exploring the parameter space, while the structured parameter editing framework keeps device complexity manageable through reusable processing templates
3Productivity
If preset processing content information is used to automate parameter editing, then the productivity of retake process is improved, but the manufacturing precision may deteriorate due to automated editing without direct user control
Solution Approach 1:
The patent implements feedback by allowing users to evaluate the generated sequence data and provide input on whether the singing manner matches their intent. This feedback loop enables the system to learn from user evaluations and refine its automated editing, maintaining high productivity while improving precision through iterative refinement based on user responses
Solution Approach 2:
The patent applies dynamics by making the processing content information adaptable and modifiable based on user feedback and evaluation results. The system dynamically adjusts its automated editing behavior, allowing the precision to improve over time while maintaining high productivity through the established automated framework
Data Source
AI summary
A voice synthesis device includes a sequence data generation unit configured to generate sequence data including a plurality of kinds of parameters for controlling vocalization of a voice to be synthesized based on music information and lyrics information, an output unit configured to output a singing voice based on the sequence data, and a processing content information acquisition unit configured to acquire a plurality of processing content information, associated with each of pieces of preset singing manner information. Each of the content information indicates contents of edit processing for all or part of the parameters. The sequence data generation unit generates a plurality of pieces of sequence data, and the sequence data are obtained by editing the all or part of the parameters included in the sequence data, based on the content information associated with one of the pieces of singing manner information specified by a user.


