Personalized Voice-Image Content Generation for Natural Digital Clones

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to create synthesized content that accurately reflects a user's voice and personal image, leading to a mismatch and awkwardness in digital presentations.

Innovation Solution

A content generation system that utilizes a voice generation model and an image generation model to learn a user's reading style and facial movements, synthesizing a digital clone that reads and gestures naturally, using machine learning to generate a personalized voice and image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a voice generation model is used to synthesize a user's voice, then the naturalness of the synthesized content is improved, but the complexity of the system increases

Engineering Contradiction:
Improvenaturalness of synthesized voiceVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The voice generation model is trained in advance using the user's reading samples to learn their unique reading style, voice characteristics, and pronunciation patterns. This preliminary training allows the model to generate natural-sounding synthesized speech without requiring complex real-time processing during actual content generation

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a digital clone by copying the user's voice characteristics, reading style, and facial expression patterns into the generation model. This copy can then reproduce the user's speech and demeanor naturally without requiring the actual user to be present, reducing system complexity while maintaining naturalness

Inventive Principle:
Principle #26Copying

2Stability of the object's composition

If a digital clone is generated to match the user's personal image, then the consistency between voice and image is improved, but the difficulty of generating accurate facial movements increases

Engineering Contradiction:
Improvevoice-image consistencyVSAvoidfacial movement accuracy
Core Design Contradiction:
Stability of the object's compositionVSDifficulty of detecting and measuring

Solution Approach 1:

The image generation model is trained using feedback from the user's actual facial movements and expressions during reading. The model learns to synchronize facial gestures with speech patterns, creating a digital clone that naturally coordinates voice and image without requiring complex real-time detection and adjustment mechanisms

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If machine learning is used to learn the user's reading style, then the personalization of the synthesized content is improved, but the amount of training data required increases

Engineering Contradiction:
Improvepersonalization of synthesized contentVSAvoidtraining data volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system achieves effective personalization by learning from a limited set of representative reading samples rather than requiring extensive training data. The voice generation model captures the essential characteristics of the user's reading style from these partial examples, sufficient to generate naturally personalized synthesized content without needing large datasets

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12608864B2Content generation device, content generation method, and program
Publication Date: 2026.04.21 TOPPAN HOLDINGS INC
  • US12608864B2 patent drawing
  • US12608864B2 patent drawing
  • US12608864B2 patent drawing

AI summary

A content generation device includes: an acquisition unit that acquires text data representing a first text, being a reading target; a voice generation unit that, using a voice generation model that based on a voice in which a user has read out a second text, being a learning target, has learned a way of reading out the second text in a voice of the user, generates a synthesized voice in which the first text represented by the acquired text data is read out in the voice of the user; and a synthesis unit that generates synthesized content by synthesizing the generated synthesized voice and a personal image of the user.