Voice Feature Conversion for Fictional Language Speech Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice generation systems require extensive dictionary data preparation to convert text from a real language to a fictional language, leading to increased developer workload as the number of character strings increases.

Innovation Solution

A voice generating system that uses a learned conversion model to convert text from a different language into voice feature amounts, synthesizing a voice that mimics a fictional language without the need for extensive dictionary data preparation, by converting characters and tokens into voice feature amounts using neural networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If dictionary data is prepared to associate converted character strings in a fictional language with all character strings that may be included in the text to be converted into voice, then a voice in a fictional language can be generated, but the work burden on a developer who prepares the dictionary data increases

Engineering Contradiction:
Improveability to generate voice in fictional languageVSAvoidwork burden on developer
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system enables self-service by automatically generating dictionary data through the conversion process itself. When converting text from a real language to a fictional language, the system automatically creates and stores the mapping between original character strings and converted character strings, eliminating the need for manual dictionary preparation by developers.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by pre-generating and storing dictionary data through automatic conversion processes before actual voice generation is needed. The dictionary data is created in advance through automated text conversion, so that when voice generation is required, the mapping data is already available without requiring manual preparation.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If the number of types of character strings that may be included in the text to be converted into voice increases, then the coverage of voice generation improves, but the work burden on a developer who prepares the dictionary data increases

Engineering Contradiction:
Improvecoverage of character stringsVSAvoidtime for dictionary preparation
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system automatically generates dictionary entries for new character strings through self-service conversion processes. When new types of character strings appear in the input text, the system automatically creates the corresponding mappings to fictional language character strings and stores them, expanding coverage without requiring additional manual work from developers.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary conversion and storage of character string mappings automatically. By pre-processing text and generating dictionary data through automated conversion, the system expands its coverage of character string types without proportionally increasing the time required for manual dictionary preparation.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12518736B2Non-transitory computer-readable medium and voice generating system
Publication Date: 2026.01.06 SQUARE ENIX HLDG CO LTD
  • US12518736B2 patent drawing
  • US12518736B2 patent drawing
  • US12518736B2 patent drawing

AI summary

According to one or more embodiments, a non-transitory computer-readable medium including a program that, when executed, causes a server to perform functions including: converting a first text into a voice feature amount by inputting the first text into a learned conversion model, wherein the first text is in a different language from a predetermined language, the learned conversion model is pre-learned to convert a second text in the predetermined language into a voice feature amount, and synthesizing a voice from the converted voice feature amount.