Adaptive Audio Mixing via Neural Network Stem Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Commercial video games rely on pre-recorded audio tracks that provide a predictable and repetitive audio experience, failing to adapt to dynamic game scenarios and user interactions.

Innovation Solution

A neural network-based adaptive audio mixing system that selects and mixes pre-recorded music stems in real-time based on game scene characteristics, user interactions, and emotional responses, using reinforcement learning to adjust parameters and create a unique, non-deterministic audio mix.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If pre-recorded audio tracks are used, then audio quality is maintained, but audio experience becomes repetitive and predictable

Engineering Contradiction:
Improveaudio qualityVSAvoidaudio experience variability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system transitions from static pre-recorded audio tracks to dynamic real-time audio mixing. A trained neural network dynamically selects and mixes audio stems based on real-time game state parameters, player actions, and contextual factors, creating continuously varying audio experiences while maintaining professional quality through controlled mixing of pre-recorded elements.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The audio track is segmented into separate stems (instrumental layers, melodic elements, rhythmic components) that can be independently selected and mixed. This segmentation allows the system to recombine these elements in numerous variations, preventing repetition while maintaining the quality of individual components.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If real-time adaptive mixing is implemented, then audio experience becomes dynamic and engaging, but system complexity increases

Engineering Contradiction:
Improveaudio experience variabilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Audio stems are pre-recorded and pre-processed into compatible segments during production. This preliminary action allows the runtime system to focus only on selection and mixing parameters rather than generating audio from scratch, significantly reducing computational complexity while maintaining adaptability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of requiring complex real-time audio synthesis capabilities, the system uses pre-recorded stems as templates or copies of professional audio material. The neural network learns to mix these copies based on game state, avoiding the need for complex generative audio models while achieving dynamic adaptability.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If multiple audio stems are mixed in real-time, then audio variety increases, but processing time and computational resources increase

Engineering Contradiction:
Improveaudio mix varietyVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system controls the number of audio stems mixed at any given time by adjusting mixing parameters such as stem selection, volume levels, panning positions, and crossfade durations. The neural network learns optimal parameter combinations for different game states, achieving variety through parameter variation rather than processing excessive numbers of stems simultaneously.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If pre-recorded tracks are used, then production efficiency is maintained, but user engagement decreases due to repetition

Engineering Contradiction:
Improveproduction efficiencyVSAvoiduser engagement
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system automatically mixes and adapts audio content based on game state without requiring manual composition for each scenario. The trained neural network serves itself by learning from training data and making autonomous mixing decisions, maintaining production efficiency while dramatically improving adaptability and user engagement through dynamic audio variation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11839815B2Adaptive audio mixing
Publication Date: 2023.12.12 ATI TECHNOLOGIES ULC
  • US11839815B2 patent drawing
  • US11839815B2 patent drawing
  • US11839815B2 patent drawing

AI summary

Systems, apparatuses, and methods for performing adaptive audio mixing are disclosed. A trained neural network dynamically selects and mixes pre-recorded, human-composed music stems that are composed as mutually compatible sets. Stem and track selection, volume mixing, filtering, dynamic compression, acoustical/reverberant characteristics, segues, tempo, beat-matching and crossfading parameters generated by the neural network are inferred from the game scene characteristics and other dynamically changing factors. The trained neural network selects an artist's pre-recorded stems and mixes the stems in real-time in unique ways to dynamically adjust and modify background music based on factors such as game scenario, the unique storyline of the player, scene elements, the player's profile, interest, and performance, adjustments made to game controls (e.g., music volume), number of viewers, received comments, player's popularity, player's native language, player's presence, and/or other factors. The trained neural network creates unique music that dynamically varies according to real-time circumstances.