ML Audio Mix Automation for Dialogue Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The current manual process of creating Music and Effects (M&E) sound mixes for foreign language versions of media content is labor-intensive, time-consuming, and prone to human error, increasing production costs and reducing the time available for creative collaboration between sound mixers and filmmakers.

Innovation Solution

A machine learning-based system that extracts metadata and content feature data from an original sound mix, uses a trained model to calculate and derive an M&E sound mix, and automatically generates foreign language sound mixes by combining the derived M&E mix with foreign language dialogue tracks, reducing manual intervention and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual process is used to create M&E sound mixes, then quality control and creative collaboration are maintained, but production time and labor costs increase significantly

Engineering Contradiction:
ImproveM&E mix creation speedVSAvoidProduction time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of M&E mix creation with an automated system using machine learning models and audio signal processing algorithms. The system automatically identifies dialogue segments, separates them from music and effects, and generates M&E mixes without manual intervention, thereby dramatically increasing productivity while reducing production time.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables the sound mix itself to be processed automatically through algorithmic dialogue detection and audio separation. The machine learning model autonomously analyzes the audio signal, identifies speech segments, and performs the separation task without requiring human operators, allowing the process to serve itself and eliminating labor-intensive manual work.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual dialogue identification and removal is performed, then accuracy in preserving non-dialogue sounds is maintained, but human error and inconsistency increase

Engineering Contradiction:
ImproveDialogue identification accuracyVSAvoidConsistency of M&E mix quality
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent replaces manual dialogue identification with automated machine learning-based speech detection algorithms. These algorithms consistently identify dialogue segments across different audio inputs with high accuracy and uniformity, eliminating the variability and potential errors associated with manual human review and decision-making.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system incorporates trained machine learning models that have been fed back from extensive training data to improve their accuracy. The models use feedback from training on labeled audio data to continuously refine their dialogue detection and separation capabilities, ensuring consistent and reliable performance across diverse audio inputs.

Inventive Principle:
Principle #23Feedback

3Ease of manufacture

If extensive manual processing is used to separate dialogue from music and effects, then detailed control over sound elements is achieved, but production costs and complexity increase

Engineering Contradiction:
ImproveM&E mix production simplicityVSAvoidProcessing system complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent replaces complex manual processing workflows with an automated machine learning system that handles dialogue separation in a single integrated process. The system uses trained models to automatically distinguish dialogue from music and effects, simplifying the manufacturing process while managing the computational complexity internally through algorithmic operations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Adaptability or versatility

If traditional methods are used to create foreign language versions, then cultural and linguistic nuances are preserved through human judgment, but time available for creative collaboration is reduced

Engineering Contradiction:
ImproveForeign language version adaptabilityVSAvoidTime for creative collaboration
Core Design Contradiction:
Adaptability or versatilityVSDuration of action of moving object

Solution Approach 1:

The patent replaces manual M&E mix creation with automated machine learning processing, dramatically reducing the time required to produce foreign language versions. This time savings can be reallocated to enhance creative collaboration between sound mixers and filmmakers, while the system maintains adaptability through configurable processing parameters and trained models that can be adjusted for different linguistic and cultural requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11087738B2System and method for music and effects sound mix creation in audio soundtrack versioning
Publication Date: 2021.08.10 LUCASFILM ENTERTAINMENT COMPANY LTD
  • US11087738B2 patent drawing
  • US11087738B2 patent drawing
  • US11087738B2 patent drawing

AI summary

Implementations of the disclosure describe systems and methods that leverage machine learning to automate the process of creating music and effects mixes from original sound mixes including domestic dialogue. In some implementations, a method includes: receiving a sound mix including human dialogue; extracting metadata from the sound mix, where the extracted metadata categorizes the sound mix; extracting content feature data from the sound mix, the extracted content feature data including an identification of the human dialogue and instances or times the human dialogue occurs within the sound mix; automatically calculating, with a trained model, content feature data of a music and effects (M&E) sound mix using at least the extracted metadata and the extracted content feature data of the sound mix; and deriving the M&E sound mix using at least the calculated content feature data.