Voice Synthesis Style Assignment via Time Range Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice synthesis systems require users to tediously re-designate the pronunciation style for each note when editing, making the process cumbersome and time-consuming.

Innovation Solution

An information processing method and device that allows users to set a pronunciation style for a specific range on a time axis, arrange notes within that range, and generate a characteristic transition of acoustic characteristics, reducing the need for repetitive style designation by using transition estimation models and machine learning to reflect underlying trends in the data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If users designate pronunciation style for each note individually, then the voice synthesis accuracy is improved, but the user workload and time consumption increase significantly

Engineering Contradiction:
Improvevoice synthesis accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the time axis into specific ranges and applies pronunciation styles to each range rather than to individual notes. This segmentation approach maintains synthesis accuracy within each range while significantly reducing the number of style designations required, thereby resolving the contradiction between accuracy and time consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by setting pronunciation styles for specific time ranges before notes are arranged. This allows the system to pre-establish acoustic characteristics and transitions, so that when notes are later arranged within those ranges, the style is already determined, eliminating the need for repeated style designations and reducing time consumption.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If users re-designate pronunciation style for each edited note, then the context accuracy is improved, but the editing process becomes cumbersome

Engineering Contradiction:
Improvecontext accuracyVSAvoidediting ease
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent merges the pronunciation style designation with the time range setting. When users edit notes within a specific range, the pronunciation style is automatically applied to the entire range rather than requiring re-designation for each note. This merging maintains context accuracy while dramatically improving editing ease.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system provides self-service by automatically maintaining and applying pronunciation styles to edited notes within a designated range. When users add, remove, or modify notes, the system automatically ensures the appropriate style is applied without requiring manual re-designation, thereby improving ease of operation while maintaining reliability.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If pronunciation style is set for entire piece, then the operation simplicity is improved, but the contextual accuracy for specific sections deteriorates

Engineering Contradiction:
Improveoperation simplicityVSAvoidcontextual accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent applies local quality by allowing different pronunciation styles to be assigned to different specific ranges within the musical piece. Users can set simple global styles for overall simplicity while also defining local ranges with specific styles for sections requiring contextual accuracy, thereby resolving the contradiction between operation simplicity and contextual accuracy.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11437016B2Information processing method, information processing device, and program
Publication Date: 2022.09.06 YAMAHA CORP
  • US11437016B2 patent drawing
  • US11437016B2 patent drawing
  • US11437016B2 patent drawing

AI summary

An information processing method is realized by a computer, and includes setting a pronunciation style with regard to a specific range on a time axis, arranging one or more notes in accordance with an instruction from a user within the specific range for which the pronunciation style has been set, and generating a characteristic transition, which is a transition of acoustic characteristics of voice that pronounces the one or more notes within the specific range in the pronunciation style set for the specific range.