Audio Mapping Using Neural Networks for Automated Channel Assignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for classifying audio tracks in audio-visual content are manually intensive, making them slow and costly, necessitating a solution for automation to reduce time and human involvement in content classification and management.

Innovation Solution

The implementation of an automated audio mapping system using an artificial neural network (ANN) that processes audio-visual content to identify and classify audio tracks, such as music, dialog, and special effects, by analyzing acoustic characteristics and assigning them to predetermined channels, thereby reducing manual intervention.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual classification methods are used for audio tracks, then classification accuracy can be maintained through human judgment, but the process becomes slow and costly

Engineering Contradiction:
Improveclassification accuracyVSAvoidclassification speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual human classification (mechanical system) with an automated audio mapping system using artificial neural networks. The system processes audio tracks through multiple neural network models that analyze acoustic characteristics and automatically classify audio components into standard channels, eliminating the need for manual human review while maintaining classification accuracy through sophisticated machine learning algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual classification methods are used for audio tracks, then complex audio analysis can be performed with human expertise, but the time and human involvement required increase significantly

Engineering Contradiction:
Improveclassification qualityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-training multiple specialized neural network models on extensive audio data before deployment. The system pre-establishes audio mapping rules and channel assignments based on acoustic characteristics. When actual classification is needed, these pre-trained models immediately process audio tracks without requiring real-time human expertise, significantly reducing processing time while maintaining reliable classification quality.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If automated methods are implemented for audio classification, then processing speed and efficiency improve, but system complexity increases

Engineering Contradiction:
Improveclassification efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the audio classification task into multiple specialized neural network models, each handling specific audio components (e.g., dialogue, music, sound effects). The system segments the classification process into distinct stages: audio feature extraction, component identification, channel mapping, and quality verification. This modular segmentation improves processing efficiency while managing system complexity through organized, specialized sub-systems.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11523186B2Automated audio mapping using an artificial neural network
Publication Date: 2022.12.06 DISNEY ENTERPRISES INC
  • US11523186B2 patent drawing
  • US11523186B2 patent drawing
  • US11523186B2 patent drawing

AI summary

According to one implementation, an automated audio mapping system includes a computing platform having a hardware processor and a system memory storing an audio mapping software code including an artificial neural network (ANN) trained to identify multiple different audio content types. The hardware processor is configured to execute the audio mapping software code to receive content including multiple audio tracks, and to identify, without using the ANN, a first music track and a second music track of the multiple audio tracks. The hardware processor is further configured to execute the audio mapping software code to identify, using the ANN, the audio content type of each of the multiple audio tracks except the first music track and the second music track, and to output a mapped content file including the multiple audio tracks each assigned to a respective one predetermined audio channel based on its identified audio content type.