Audio Generation System Using Latent Space Parameterization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The current process of generating and modifying spoken audio content is time-consuming and costly, especially when localizing content for different languages or creating immersive and open-world content that requires a large range of voice lines.

Innovation Solution

The use of an autoencoder with a latent space representation to efficiently generate and modify voice content, allowing for real-time applications, personalized experiences, and improved localization by parameterizing sample voices and generating voice characteristics based on emotional states, age, nationality, and other qualities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional voice generation methods are used with actors recording lines, then voice quality and emotional expression are improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvevoice qualityVSAvoidtime consumption
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses autoencoders to create latent space representations that capture the essential characteristics of voice recordings. These latent representations serve as compressed copies that can be efficiently manipulated and decoded to generate new voice content without requiring repeated actor recordings, thus maintaining voice quality while reducing time consumption.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables control over voice characteristics by manipulating parameters in the latent space. By adjusting latent vectors corresponding to different attributes (emotion, age, accent), the system can generate varied voice outputs from a single recording session, eliminating the need for multiple recording sessions while maintaining emotional expression quality.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If voice content is translated into multiple languages for localization, then content accessibility is improved, but time and effort investment increases

Engineering Contradiction:
Improvelanguage coverageVSAvoidlocalization time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent creates a universal latent space representation that can be decoded into multiple languages. A single voice recording is transformed into a language-agnostic latent representation, which can then be synthesized in different languages by modifying the text input while keeping the voice characteristics consistent, enabling efficient multi-language localization without separate recording sessions for each language.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If pre-generated voice content is used to populate interactive environments, then content completeness is improved, but flexibility and personalization are reduced

Engineering Contradiction:
Improvevoice content volumeVSAvoidpersonalization capability
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The system transitions from static pre-generated voice content to dynamic on-the-fly synthesis. Voice content is generated in real-time by decoding latent representations with different text inputs and parameter adjustments, allowing the same voice model to adapt to different contexts, emotions, and user preferences, thus providing both content volume and personalization flexibility.

Inventive Principle:
Principle #15Dynamics

4Reliability

If more voice lines are recorded to reduce repetition in open-world content, then content immersion is improved, but recording time and cost increase

Engineering Contradiction:
Improvecontent immersionVSAvoidrecording time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent enables generation of diverse voice lines by manipulating parameters in the latent space such as emotion, pitch, and timing. A single recorded voice can be transformed into numerous variations by adjusting these parameters programmatically, creating the appearance of extensive voice content for immersive environments without requiring proportional recording time.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250131934A1Audio generation system and method
Publication Date: 2025.04.24 SONY INTERACTIVE ENTERTAINMENT LLC
  • US20250131934A1 patent drawing
  • US20250131934A1 patent drawing
  • US20250131934A1 patent drawing

AI summary

An audio generation system for generating output audio comprising speech, the system comprising an input unit configured to receive a first input defining the semantic content of the output audio, and a second input defining one or more desired characteristics of the output audio, a parameter identification unit configured to identify, from one or more latent spaces each associated with one or more possible characteristics of the output audio, one or more parameters for use in generating the output audio in dependence upon the second input, and an output generating unit configured to generate output audio in dependence upon the first input and the identified one or more parameters.