Automated Voice Casting Using Neural Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The voice casting process for media localization is laborious and subjective, lacking objective metrics and well-defined vocabulary for voice similarity, making it difficult to quantify and repeat the selection of voice actors that closely match the original actors in a different language.
Innovation Solution
An automated system that uses deep learning neural networks to analyze voice recordings, generate multi-dimensional encodings, and compare similarity scores for candidate voice actors, allowing for objective and repeatable voice casting by identifying and combining utterance types to create balanced voice samples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual voice casting is used with human voice casters listening to voice clips, then subjective voice similarity assessment can be performed, but the process is laborious and lacks objective metrics for quantification
Solution Approach 1:
The patent replaces the manual mechanical process of human voice casters listening to and comparing voice clips with an automated computer-based system that uses signal processing and machine learning algorithms to objectively measure voice similarity. The system extracts acoustic features from voice samples, computes similarity metrics, and ranks candidate voice actors automatically, eliminating the laborious manual comparison process while providing quantifiable objective measurements of voice similarity.
2Loss of time
If limited voice clips are provided in Voice Testing Kits, then the voice casting process can be completed in reasonable time, but it is difficult to obtain additional voice samples when the voice caster is undecided
Solution Approach 1:
The patent applies preliminary action by automatically extracting and organizing multiple voice samples from available media content before the voice casting decision is made. The system pre-processes large amounts of audio data, extracts relevant utterances, and prepares comprehensive voice samples for analysis, so that when the system needs to evaluate additional samples, they are already prepared and available for immediate comparison without extending the decision-making time.
3Adaptability or versatility
If multiple voice casters perform subjective assessment, then different perspectives can be considered, but the selections vary without clear explanation of reasoning
Solution Approach 1:
The patent implements feedback by providing the automated system with detailed quantitative measurements of voice similarity, including specific acoustic feature comparisons and similarity scores. The system feedback includes breakdowns of which acoustic characteristics contribute most to the similarity assessment, providing clear explanatory reasoning for each voice actor ranking that replaces the unexplained subjective judgments of human voice casters.
4Quantity of substance
If brief voice samples of around 60 seconds are used in Voice Testing Kits, then the kits remain manageable in size, but insufficient data is available for accurate voice matching
Solution Approach 1:
The patent applies segmentation by automatically dividing available media content into multiple distinct voice samples, each containing specific utterances or phrases. The system segments the audio data into manageable units with associated metadata, organizing them in a structured format that allows comprehensive analysis while keeping the overall system complexity manageable through automated processing and systematic organization of the segmented voice data.
Data Source
AI summary
A method and system for automated voice casting compares candidate voices samples from candidate speakers in a target language with a primary voice sample from a primary speaker in a primary language. Utterances in the audio samples of the candidates speakers and the primary speaker are identified and typed and voice samples generated that meet applicable utterance type criteria. A neural network is used to generate an embedding for the voice samples. A voice sample can include groups of different utterance types and embeddings generated for each utterance group in the voice sample and then combined in a weighted form wherein the resulting embedding emphasizes selected utterance types. Similarities between embeddings for the candidate voice samples relative to the primary voice sample are evaluated and used to select a candidate speaker that is a vocal match.


