Emotional Speech Synthesis Using Face Feature Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis techniques do not account for speaker-specific factors such as age, height, weight, and emotional expression, resulting in synthesized speech that lacks personalization and emotional depth.

Innovation Solution

A speech processing apparatus that separates moving image data into frames and extracts face feature points, generates a first generation network to produce face feature points from speech data, evaluates the network's appropriateness, and then generates a second generation network to synthesize emotional speech based on designated fixed and uncertain settings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing speech synthesis techniques are used, then speech can be synthesized with basic quality, but speaker-specific factors such as age, height, weight, and emotional expression are not accounted for

Engineering Contradiction:
Improvepersonalization capabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the speech synthesis process into multiple independent modules: a first generation network for generating face feature points from speech data, and a second generation network for synthesizing speech with emotional expression. This segmentation allows each network to specialize in specific aspects, improving personalization capability while maintaining manageable system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a nested structure where the first generation network is integrated within the second generation network system. The first network processes speech data to generate face feature points, which then serve as input to the second network for final speech synthesis. This nesting allows the system to handle multiple processing stages efficiently, enabling personalization based on speaker-specific factors without proportionally increasing overall system complexity.

Inventive Principle:
Principle #7Nested doll (Nesting)

2Manufacturing precision

If emotional expression is added to speech synthesis, then speech quality and naturalness improve, but the complexity of processing and controlling emotions increases

Engineering Contradiction:
Improvespeech qualityVSAvoidemotion processing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent extracts emotional expression control from the main speech synthesis process by using a separate first generation network that specifically processes emotional features. This network generates face feature points that capture emotional information, which are then used by the second network for speech synthesis. By taking out emotion processing into a separate module, the system improves speech quality while managing emotion processing complexity through isolation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The first generation network acts as an intermediary between the speech input and the final speech synthesis output. It processes speech data to generate face feature points that encode emotional information, serving as a mediator that translates raw speech data into emotionally enriched features for the second network, thereby improving speech quality without directly increasing the complexity of the main synthesis process.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If multiple settings (fixed and uncertain) are considered for speech synthesis, then personalization and emotional accuracy improve, but the time required for processing and network generation increases

Engineering Contradiction:
Improvespeaker setting accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing by pre-separating moving image data into frames and pre-extracting face feature point data before the main synthesis process. This preliminary action prepares the data in advance, reducing the computational burden during actual speech synthesis and thereby improving speaker setting accuracy without proportionally increasing processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic processing where the first generation network adapts to different speaker settings and emotional states in real-time. The network dynamically adjusts face feature point generation based on the specific combination of fixed settings (age, height, weight) and uncertain settings (emotions, utterance content), enabling precise personalization while maintaining efficient processing through adaptive algorithms.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3693957B1Voice processing device and program
Publication Date: 2025.04.23 KAINUMA KEN ICHI
  • EP3693957B1 patent drawingFigure 1
  • EP3693957B1 patent drawingFigure 2
  • EP3693957B1 patent drawingFigure 3~4

AI summary

Synthesis of emotional speech is realized while settings unique to each speaker is taken into account. A speech processing apparatus is provided in which, while face feature points are extracted from moving image data obtained by imaging a face of a speaker, for each frame, a first generation network for generating face feature points of the corresponding frame on the basis of speech feature data extracted from uttered speech of the speaker for each frame is generated, and whether or not the first generation network is appropriate is evaluated using an identification network, then, a second generation network for generating the uttered speech from a plurality of uncertain settings including at least text representing utterance content of the uttered speech and information indicating emotions included in the uttered speech, a plurality of types of fixed settings which define speech quality of the speaker, and the face feature points generated by the first generation network evaluated as appropriate, is generated, and whether or not the second generation network is appropriate is evaluated using the above-described identification network.