Identity Transfer Models for Audio Content Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating audio and video content are time-consuming and face challenges in remixing media by combining aspects from multiple sources, such as merging a song with a different artist's voice, due to limitations in extracting and integrating intertwined media elements.

Innovation Solution

The implementation of machine-learning identity transfer models that allow users to select audio or video content and a target identity, processing it to generate synthesized media by mapping content embeddings with voice or visual identity embeddings, enabling the creation of remixed content without requiring the actual recording of the target artist.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional recording or filming methods are used to generate audio and video content, then the content can be created with authentic artist performance, but the process is time-consuming and requires actual recording sessions

Engineering Contradiction:
Improveauthenticity of artist performanceVSAvoidtime required for recording or filming
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses machine learning models to create synthetic copies of artist voices and appearances. The system extracts identity features from reference recordings and generates new audio content that mimics the target artist's voice characteristics without requiring the artist to actually record the new content. This allows rapid generation of authentic-sounding content while eliminating the time-consuming recording process.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical recording process with a computational system. Instead of physically recording an artist's performance in a studio, the system uses neural networks and identity transfer models to synthesize the performance. This substitution of mechanical recording with computational generation dramatically reduces time while maintaining authenticity through learned identity characteristics.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If machine-learning techniques are used to generate media content, then the process can be faster and more efficient, but there are limitations in extracting and combining intertwined aspects from multiple sources

Engineering Contradiction:
Improvespeed of content generationVSAvoidcomplexity of extracting and integrating media elements
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the media content into distinct identity-related components that can be independently processed. The system separates voice identity characteristics from the underlying audio content, allowing these elements to be extracted and recombined flexibly. This segmentation enables the system to handle intertwined aspects by treating them as separable identity features rather than inseparable mixed signals.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces identity embeddings as an intermediary representation that bridges multiple sources. The neural network models convert diverse input data (audio recordings, visual content) into standardized identity embedding vectors, which then serve as intermediaries for combining aspects from different sources. This intermediary representation simplifies the integration process by providing a common framework for merging intertwined elements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12087268B1Identity transfer models for generating audio/video content
Publication Date: 2024.09.10 AMAZON TECH INC
  • US12087268B1 patent drawing
  • US12087268B1 patent drawing
  • US12087268B1 patent drawing

AI summary

Systems, devices, and methods are provided for training and/or inferencing using machine-learning models. In at least one embodiment, a user selects a source media (e.g., video or audio file) and a target identity. A content embedding may be extracted from the source media, and an identity embedding may be obtained for the target identity. The content embedding of the source media and the identity embedding of the target identity may be provided to a transfer model that generates synthesized media. For example, a user may select a song that is sung by a first artist and then select a second artist as the target identity to produce a cover of the song in the voice of the second artist.