Voice Synthesis Parameter Synchronization with Image Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice synthesis technologies fail to maintain a balanced impression between synthesized voices and accompanying images, leading to an undesirable imbalance when voice synthesis parameters are changed.

Innovation Solution

An information processing method and device that synchronizes changes in voice parameters with corresponding changes in image parameters, ensuring that both voice and image are adjusted in real-time, preventing imbalance during playback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If voice synthesis parameters are changed to improve voice quality, then voice synthesis capability is improved, but balance between voice and accompanying image deteriorates

Engineering Contradiction:
Improvevoice synthesis qualityVSAvoidbalance between voice and image
Core Design Contradiction:
Manufacturing precisionVSStability of the object's composition

Solution Approach 1:

The patent merges the control of voice parameters and image parameters into a single integrated system. When a user adjusts voice synthesis parameters through the UI, the system automatically adjusts corresponding image parameters to maintain balance. This combining of parameter controls resolves the contradiction by ensuring that voice quality improvement does not occur at the expense of voice-image balance.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements coordinated parameter changes across multiple domains. When voice parameters (such as pitch, timbre, or volume) are modified, the system automatically changes corresponding image parameters (such as character expression, mouth movement, or visual intensity) to maintain harmonious balance. This coordinated parameter adjustment resolves the technical contradiction between improving voice synthesis quality and maintaining voice-image balance.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If only voice parameters are adjusted to improve voice synthesis, then voice parameter control is simplified, but overall content quality deteriorates due to image-voice imbalance

Engineering Contradiction:
Improvevoice parameter adjustmentVSAvoidcontent quality balance
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The system implements self-service automation where the image parameter adjustment occurs automatically without requiring separate user input. When users adjust voice parameters, the system autonomously determines and applies appropriate image parameter changes based on pre-established correspondence relationships. This self-service mechanism maintains ease of operation while ensuring overall content quality balance.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system establishes a feedback loop where voice parameter changes trigger automatic image parameter adjustments. The correspondence relationships between voice and image parameters create a feedback mechanism that continuously maintains balance. This feedback-based approach allows simplified user operation while preserving high content quality through automatic coordination.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9997153B2Information processing method and information processing device
Publication Date: 2018.06.12 YAMAHA CORP
  • US9997153B2 patent drawing
  • US9997153B2 patent drawing
  • US9997153B2 patent drawing

AI summary

An information processing method includes receiving a change instruction to change a voice parameter used in synthesizing a voice for a set of texts, changing the voice parameter in accordance with the change instruction to change the voice parameter, changing, in accordance with the change instruction, an image parameter used in synthesizing an image of a virtual object, the virtual object indicating a character that vocalizes the voice that has been synthesized, synthesizing the voice using the changed voice parameter, and synthesizing the image using the changed image parameter.