Automatic Surround Mixer Using Rule-Based Stem Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The creation of surround audio mixes for multi-channel reproduction is costly and requires additional mixing sessions with artists and engineers, making it unaffordable for content owners and record companies, and existing technologies lack efficient methods for automatically generating high-quality surround mixes from stereo artistic mixes.
Innovation Solution
An automatic surround mixer system that processes stereo artistic mix stems and metadata to generate a surround audio mix using a mixing matrix and rule engine, which applies level, spectrum, and dynamic range modifications based on metadata and user preferences, allowing for the creation of customized surround mixes without operator participation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual mixing sessions are conducted to create surround audio mixes, then the quality and artistic value of the mix is improved, but the cost and time required increases significantly
Solution Approach 1:
The system creates a virtual copy of the mixing engineer's expertise through AI models trained on professional mixing data. The virtual mixing engineer processes stems to generate surround mixes that replicate the quality of manual mixing without requiring actual human mixing sessions, thereby resolving the contradiction between mix quality and time consumption
Solution Approach 2:
The system enables automatic self-mixing of audio stems through AI-powered virtual mixing engineers. The process autonomously analyzes stems, applies appropriate processing, and generates surround mixes without human intervention, eliminating the need for manual mixing sessions while maintaining professional quality standards
2Manufacturing precision
If additional mixing sessions are conducted for surround audio, then the surround mix quality is improved, but the cost becomes unaffordable for content owners
Solution Approach 1:
The patent creates virtual copies of expert mixing capabilities through AI models trained on professional surround mixing data. These virtual mixing engineers provide expert-level surround mix generation at a fraction of the cost of actual mixing sessions, making high-quality surround mixes affordable for content owners while maintaining professional quality standards
Solution Approach 2:
The system changes the fundamental parameter of how mixing is performed by transitioning from human-operated mixing sessions to AI-based automatic mixing. This parameter change dramatically reduces cost while maintaining quality through sophisticated algorithms that analyze stems and apply appropriate processing based on learned professional mixing patterns
3Productivity
If automatic mixing is implemented, then the cost and time efficiency is improved, but the ability to capture artistic intent deteriorates
Solution Approach 1:
The system copies artistic intent by training AI models on extensive datasets of professional mixing work. The virtual mixing engineers learn to recognize and replicate artistic decisions, genre-specific conventions, and creative preferences from the training data, enabling automatic mixing that preserves artistic intent while dramatically improving efficiency
Solution Approach 2:
The system incorporates feedback mechanisms where the AI models are trained on professional mixing outcomes and continuously refine their understanding of artistic intent. By learning from feedback in the training data and through iterative improvement, the virtual mixing engineers capture nuanced artistic decisions while maintaining high productivity
Data Source
Figure 1
Figure 2A~2C
Figure 3
AI summary
There are disclosed automatic mixers and methods for creating a surround audio mix. A set of rules may be stored in a rule base. A rule engine may select a subset of the set of rules based, at least in part, on metadata associated with a plurality of stems. A mixing matrix may mix the plurality of stems in accordance with the selected subset of rules to provide three or more output channels.