Voice Synthesizing Device Abstract Quality Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice synthesis technologies struggle to generate synthetic sounds with desired voice qualities, as they primarily rely on direct parameter changes or specific characteristics, failing to effectively express abstract voice qualities like 'cute' or 'fresh' voices.

Innovation Solution

A voice synthesizing device that includes a score transformation model to convert upper level, abstract voice quality expressions into lower level, more concrete expressions, allowing users to specify and edit voice qualities using a multistage transformation process, enabling the generation of synthetic sounds with desired voice qualities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional voice synthesis technologies directly change parameters of acoustic models or reflect specified characteristics, then the control over synthetic sound is flexible and various types of synthetic sounds can be generated, but the ability to express abstract voice qualities (such as 'cute' or 'fresh' voices) is insufficient

Engineering Contradiction:
Improveability to express abstract voice qualitiesVSAvoidintuitiveness of voice quality specification
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent segments the voice quality specification process into multiple hierarchical levels. Abstract voice qualities (upper level) are divided into intermediate characteristics, which are further divided into concrete acoustic parameters (lower level). This segmentation allows users to specify abstract qualities without directly manipulating complex parameters, resolving the contradiction between versatility and ease of operation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate voice quality characteristics as mediators between abstract user specifications and concrete acoustic parameters. These intermediate characteristics serve as a bridge, translating abstract concepts like 'cute' or 'fresh' into controllable parameter adjustments, thereby enabling both abstract expression and intuitive control.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If users specify voice quality based on abstract words, then the voice quality editing becomes more intuitive and user-friendly, but the direct control over acoustic model parameters is lost

Engineering Contradiction:
Improveintuitiveness of voice quality specificationVSAvoidprecision of voice quality control
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The hierarchical segmentation maintains precision by breaking down abstract specifications into progressively more concrete levels. Each level of segmentation preserves control precision while improving intuitiveness, as the system translates high-level abstract specifications through intermediate levels to precise parameter adjustments at the lowest level.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Intermediate voice quality characteristics act as mediators that preserve control precision while enabling intuitive specification. These intermediates maintain the relationship between abstract user intent and concrete parameter values, ensuring that abstract specifications are accurately translated into precise acoustic parameter adjustments without losing control precision.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10535335B2Voice synthesizing device, voice synthesizing method, and computer program product
Publication Date: 2020.01.14 TOSHIBA DIGITAL SOLUTIONS CORP
  • US10535335B2 patent drawing
  • US10535335B2 patent drawing
  • US10535335B2 patent drawing

AI summary

According to one embodiment, a voice synthesizing device includes a first operation receiving unit, a score transforming unit, and a voice synthesizing unit. The first operation receiving unit configured to receive a first operation specifying voice quality of a desired voice based on one or more upper level expressions indicating the voice quality. The score transforming unit configured to transform, based on a score transformation model that transforms a score of the upper level expression into a score of a lower level expression which is less abstract than the upper level expression, the score of the upper level expression corresponding to the first operation into a score of one or more lower level expressions. The voice synthesizing unit configured to generate a synthetic sound corresponding to a certain text based on the score of the lower level expression.