Voice Synthesis Style Assignment via Time Range Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice synthesis systems require users to tediously re-designate the pronunciation style for each note when editing, making the process cumbersome and time-consuming.
Innovation Solution
An information processing method and device that allows users to set a pronunciation style for a specific range on a time axis, arrange notes within that range, and generate a characteristic transition of acoustic characteristics, reducing the need for repetitive style designation by using transition estimation models and machine learning to reflect underlying trends in the data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users designate pronunciation style for each note individually, then the voice synthesis accuracy is improved, but the user workload and time consumption increase significantly
Solution Approach 1:
The patent segments the time axis into specific ranges and applies pronunciation styles to each range rather than to individual notes. This segmentation approach maintains synthesis accuracy within each range while significantly reducing the number of style designations required, thereby resolving the contradiction between accuracy and time consumption.
Solution Approach 2:
The patent performs preliminary action by setting pronunciation styles for specific time ranges before notes are arranged. This allows the system to pre-establish acoustic characteristics and transitions, so that when notes are later arranged within those ranges, the style is already determined, eliminating the need for repeated style designations and reducing time consumption.
2Reliability
If users re-designate pronunciation style for each edited note, then the context accuracy is improved, but the editing process becomes cumbersome
Solution Approach 1:
The patent merges the pronunciation style designation with the time range setting. When users edit notes within a specific range, the pronunciation style is automatically applied to the entire range rather than requiring re-designation for each note. This merging maintains context accuracy while dramatically improving editing ease.
Solution Approach 2:
The system provides self-service by automatically maintaining and applying pronunciation styles to edited notes within a designated range. When users add, remove, or modify notes, the system automatically ensures the appropriate style is applied without requiring manual re-designation, thereby improving ease of operation while maintaining reliability.
3Ease of operation
If pronunciation style is set for entire piece, then the operation simplicity is improved, but the contextual accuracy for specific sections deteriorates
Solution Approach 1:
The patent applies local quality by allowing different pronunciation styles to be assigned to different specific ranges within the musical piece. Users can set simple global styles for overall simplicity while also defining local ranges with specific styles for sections requiring contextual accuracy, thereby resolving the contradiction between operation simplicity and contextual accuracy.
Data Source
AI summary
An information processing method is realized by a computer, and includes setting a pronunciation style with regard to a specific range on a time axis, arranging one or more notes in accordance with an instruction from a user within the specific range for which the pronunciation style has been set, and generating a characteristic transition, which is a transition of acoustic characteristics of voice that pronounces the one or more notes within the specific range in the pronunciation style set for the specific range.


