Identity Transfer Models for Audio Content Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating audio and video content are time-consuming and face challenges in remixing media by combining aspects from multiple sources, such as merging a song with a different artist's voice, due to limitations in extracting and integrating intertwined media elements.
Innovation Solution
The implementation of machine-learning identity transfer models that allow users to select audio or video content and a target identity, processing it to generate synthesized media by mapping content embeddings with voice or visual identity embeddings, enabling the creation of remixed content without requiring the actual recording of the target artist.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional recording or filming methods are used to generate audio and video content, then the content can be created with authentic artist performance, but the process is time-consuming and requires actual recording sessions
Solution Approach 1:
The patent uses machine learning models to create synthetic copies of artist voices and appearances. The system extracts identity features from reference recordings and generates new audio content that mimics the target artist's voice characteristics without requiring the artist to actually record the new content. This allows rapid generation of authentic-sounding content while eliminating the time-consuming recording process.
Solution Approach 2:
The patent replaces the mechanical recording process with a computational system. Instead of physically recording an artist's performance in a studio, the system uses neural networks and identity transfer models to synthesize the performance. This substitution of mechanical recording with computational generation dramatically reduces time while maintaining authenticity through learned identity characteristics.
2Productivity
If machine-learning techniques are used to generate media content, then the process can be faster and more efficient, but there are limitations in extracting and combining intertwined aspects from multiple sources
Solution Approach 1:
The patent segments the media content into distinct identity-related components that can be independently processed. The system separates voice identity characteristics from the underlying audio content, allowing these elements to be extracted and recombined flexibly. This segmentation enables the system to handle intertwined aspects by treating them as separable identity features rather than inseparable mixed signals.
Solution Approach 2:
The patent introduces identity embeddings as an intermediary representation that bridges multiple sources. The neural network models convert diverse input data (audio recordings, visual content) into standardized identity embedding vectors, which then serve as intermediaries for combining aspects from different sources. This intermediary representation simplifies the integration process by providing a common framework for merging intertwined elements.
Data Source
AI summary
Systems, devices, and methods are provided for training and/or inferencing using machine-learning models. In at least one embodiment, a user selects a source media (e.g., video or audio file) and a target identity. A content embedding may be extracted from the source media, and an identity embedding may be obtained for the target identity. The content embedding of the source media and the identity embedding of the target identity may be provided to a transfer model that generates synthesized media. For example, a user may select a song that is sung by a first artist and then select a second artist as the target identity to produce a cover of the song in the voice of the second artist.


