A
system, apparatus and method for performing adaptive
audio mixing are disclosed. A trained neural network dynamically selects and mixes pre-recorded human-composed music stems arranged in a mutually compatible set. Stem and track selection, volume mixing, filtering, dynamic compression, acoustic /
reverberation characteristics, segue, tempo, beat matching and cross-fade parameters generated by the neural network are inferred from game scene characteristics and other dynamically changing factors. The trained neural network selects pre-recorded stems of artists and mixes the stems in a unique way in real time to dynamically adjust and change background music based on factors such as the game
scenario, the player's unique storyline, scene elements, player's profile, interests, performance, adjustments made to game controls (e.g., music volume), number of viewers, comments received, player popularity, player's native language, player's influence and / or other factors. The trained neural network generates unique music that dynamically changes according to real-time conditions.