Demixing and remixing of composite audio programs for playback in the venue

By decomposing and reconstructing audio programs to mitigate comb filtering and latency, the system provides synchronized and immersive audio experiences in venues.

JP2026508363APending Publication Date: 2026-03-10SPHERE ENTERTAINMENT GROUP LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-01
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Comb filtering effects and audio latency in real-world venues cause undesirable audio phenomena such as metallic sounds, robotic tones, and synchronization issues between multiple audio sources, which degrade the audio experience.

Method used

Systems and methods that decompose composite audio programs into multiple audio voices, analyze their characteristics, and intelligently reconstruct the audio presentation to mitigate comb filtering and latency, using audio demixing and remixing tools to ensure seamless and immersive playback.

Benefits of technology

The solution effectively reduces comb filtering and audio latency, ensuring synchronized and aesthetically appealing audio reproduction in real-world venues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026508363000001_ABST
    Figure 2026508363000001_ABST
Patent Text Reader

Abstract

Systems, methods, and devices can play a composite audio program associated with an event being hosted by a real-world venue. These systems, methods, and devices can seamlessly decompose the composite audio program into multiple audio voices that can be collectively played by the real-world venue. As part of this decomposition, these systems, methods, and devices can analyze the audio voices and identify one or more characteristics, parameters, and / or attributes of the audio voices. These systems, methods, and devices can intelligently construct an audio presentation from the multiple audio voices and play the composite audio program within the real-world venue.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 449,797, filed March 3, 2023, which is incorporated herein by reference in its entirety.

[0002] The comb filtering effect describes a phenomenon in real-world venues that occurs when an audio sound and its reflection arrive at a location within the venue at different times. The audio sound may be reflected when it comes into contact with a hard surface within the real-world venue, such as a floor, wall, window, or the like. Because the reflection travels a longer distance than the audio sound, the reflection may arrive at the location later than the audio sound. Often, certain frequencies of the audio sound are amplified or attenuated by the superposition of their reflections on themselves, causing a comb filtering effect. This superposition can cause some cancellation and amplification of the audio spectrum, which can result in a subjectively metallic-like audio sound, also referred to as metallic ringing. Generally, when the time delay between the audio sound and its reflection is approximately twenty (20) milliseconds to thirty (30) milliseconds, the human ear may undesirably perceive the audio sound and its reflection as separate signals. For example, when the time delay between an audio signal and its reflection is approximately fifty (50) milliseconds, the human ear begins to perceive the reflection as an echo of the audio signal, and this is even more apparent at approximately one hundred (100) milliseconds. However, comb filtering effects can also be present when the time delay between an audio signal and its reflection is approximately twelve (12) milliseconds to fifteen (15) milliseconds. For example, comb filtering effects can change the timbre of an audio signal at approximately a twelve (12) millisecond delay between the audio signal and its reflection. Comb filtering effects can also make an audio signal sound more "robotic" by increasing this delay between the audio signal and its reflection. In some situations, these comb filtering effects are not noticeable when the time delay between an audio signal and its reflection is approximately three (3) milliseconds to ten (10) milliseconds.

[0003] Audio latency refers to the time delay between when an audio sound is produced and when it arrives at a location within a real-world venue. In the context of a real-world venue, audio latency can be affected by various factors, including the distance between the audio source and that location within the real-world venue, the acoustics of the real-world venue, the processing and transmission of the audio sound, and the performance of any digital processing or effects on the audio sound. In many cases, multiple audio sounds generated by multiple real-world loudspeakers within a real-world venue may arrive at a single location within the real-world venue at different times. This can cause audio latency between these audio sounds within the real-world venue. In many cases, when the audio latency between the multiple audio sounds is less than about ten (10) milliseconds to about twelve (12) milliseconds, the time delay between the multiple audio sounds will more likely not be noticeable. However, when the audio latency between the multiple audio sounds is about twenty (20) milliseconds to about thirty (30) milliseconds, the human ear may perceive these audio sounds as separate signals. For example, audio latency can cause fading, reverberation, or even a lack of synchronization between multiple audio voices. In a live performance or audio production, managing audio latency between multiple audio voices is important to maintain consistent, synchronized audio. Summary of the Invention [Means for solving the problem]

[0004] Detailed Description The following disclosure provides many different embodiments or examples for implementing different features of the provided subject matter. Specific examples of components and arrangements are described herein to simplify the disclosure. These are, of course, examples only and are not intended to be limiting. Aspects of the disclosure are best understood from the following detailed description when read in conjunction with the accompanying figures. The disclosure may repeat reference numerals and / or letters in the various examples. This repetition does not, in itself, dictate a relationship between the various embodiments and / or configurations discussed. Note that, in accordance with standard practice in the industry, features have not been drawn to scale. In fact, the dimensions of features may be arbitrarily increased or decreased for clarity of discussion.

[0005] The following disclosure may include spatially relative terms such as "below," "below," "lower," "upper," "above," "upper," and the like for ease of description to describe relationships between elements or features as illustrated in the figures herein. These spatially relative terms are intended to encompass different orientations for different embodiments or examples depicted in the figures. Different embodiments or examples may be oriented differently (rotated 90 degrees or at other orientations), and the spatially relative terms contained herein may be interpreted accordingly. The following disclosure may also include the terms "about," "approximately," or "substantially" to indicate that the value of a given quantity may vary based on particular technology. Based on this technology, the terms "about" or "substantially" may indicate that the value of a given quantity varies, for example, within 1 to 15% of the value (e.g., ±1%, ±2%, ±5%, ±10%, or ±15% of the value). (Overview)

[0006] Systems, methods, and devices can play a composite audio program associated with an event being hosted by a real-world venue. These systems, methods, and devices can seamlessly decompose the composite audio program into multiple audio voices that can be collectively played by the real-world venue. As part of this decomposition, these systems, methods, and devices can analyze the audio voices and identify one or more characteristics, parameters, and / or attributes of the audio voices. These systems, methods, and devices can intelligently construct an audio presentation from the multiple audio voices and play the composite audio program within the real-world venue. As part of this construction, these systems, methods, and devices construct the audio presentation based on one or more characteristics, parameters, and / or attributes of the audio voices. After constructing the audio presentation, these systems, methods, and devices can configure the real-world venue as defined in the audio presentation to play the composite audio program within the real-world venue. As part of this construction, these systems, methods, and devices can identify audio control signals that configure the real-world venue to play the audio presentation through real-world loudspeakers within the real-world venue. [Brief explanation of the drawings]

[0007] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate the present disclosure and, together with the description, serve to further explain its principles and to enable one skilled in the art to make and use it.

[0008] [Figure 1] FIG. 1 illustrates a high-level diagrammatic representation of an example audio system that may be utilized by an example real-world venue, according to some example embodiments of the present disclosure.

[0009] [Figure 2]FIG. 2 diagrammatically illustrates the operation of an example audio demixing tool within an example audio system for decomposing a composite audio program, according to some example embodiments of the present disclosure.

[0010] [Figure 3] FIG. 3 diagrammatically illustrates the operation of an exemplary audio demixing tool for analyzing a composite audio program, according to some exemplary embodiments of the present disclosure.

[0011] [Figure 4] FIG. 4 illustrates a flowchart of an exemplary audio demixing tool according to some exemplary embodiments of the present disclosure.

[0012] [Figure 5] FIG. 5 graphically illustrates an example ensemble bounding volume that may be generated by an example audio remixing tool in an example audio system, according to some example embodiments of the present disclosure.

[0013] [Figure 6] FIG. 6 diagrammatically illustrates an example virtual venue that may be accessed by an example audio remixing tool, according to some example embodiments of the present disclosure.

[0014] [Figure 7A] 7A-7F diagrammatically illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation for playback of a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 7B] 7A-7F diagrammatically illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation for playback of a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 7C] 7A-7F diagrammatically illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation for playback of a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 7D] 7A-7F diagrammatically illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation for playback of a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 7E] 7A-7F diagrammatically illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation for playback of a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 7F] 7A-7F diagrammatically illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation for playback of a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure.

[0015] [Figure 8A] 8A-8B diagrammatically illustrate the operation of an exemplary audio remixing tool in constructing an exemplary dynamic audio presentation for playing a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 8B] 8A-8B diagrammatically illustrate the operation of an exemplary audio remixing tool in constructing an exemplary dynamic audio presentation for playing a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure.

[0016] [Figure 9] FIG. 9 graphically illustrates a high-level diagrammatic mapping of an exemplary virtual venue to an exemplary real-world venue, according to some exemplary embodiments of the present disclosure.

[0017] [Figure 10] FIG. 10 illustrates a flowchart of an exemplary audio remixing tool according to some exemplary embodiments of the present disclosure.

[0018] [Figure 11] FIG. 11 diagrammatically illustrates a simplified block diagram of a computing device that may be utilized to implement electronic devices within an exemplary real-world venue, according to some embodiments of the present disclosure.

[0019] In the accompanying drawings, like reference numbers indicate identical or functionally similar elements. Additionally, the leftmost digit(s) of a reference number identifies the drawing in which the reference number first appears. DETAILED DESCRIPTION OF THE INVENTION

[0020] Example Audio System for Use in an Example Real-World Venue FIG. 1 illustrates a high-level diagrammatic representation of an exemplary audio system that may be utilized by an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. In the exemplary embodiment illustrated in FIG. 1, audio playback system 100 can play a composite audio program associated with an event being held by real-world venue 102. For example, real-world venue 102 can represent a music real-world venue, e.g., a music theater, music club, and / or concert hall; a sports real-world venue, e.g., an arena, convention center, and / or stadium; and / or any other suitable real-world venue that would be apparent to one skilled in the art without departing from the spirit and scope of the present disclosure. As another example, the event can represent a music event, a theatrical event, a sporting event, a video, and / or any other suitable event that would be apparent to one skilled in the art without departing from the spirit and scope of the present disclosure. In some embodiments, audio playback system 100 can access a composite audio program. As described herein, audio playback system 100 can execute audio demixing tools to seamlessly decompose a composite audio program into multiple audio voices that can be collectively played back by real-world venue 102. Also, as described herein, audio playback system 100 can execute audio remixing tools to intelligently construct an audio presentation from these multiple audio voices and play back the composite audio program within real-world venue 102. In some embodiments, audio playback system 100 can include an audio playback server 104 to perform audio demixing and / or audio remixing.

[0021] 1 , the audio playback server 104, an exemplary embodiment of which is described in further detail below, can execute an audio demixing tool 152 to seamlessly decompose the composite audio program 150 into multiple audio voices that can be collectively played by the real-world venue 102. Alternatively, or in addition, the audio playback server 104 can execute an audio remixing tool 154 to intelligently construct an audio presentation from the multiple audio voices, play these audio voices within the real-world venue 102, and play the composite audio program 150 within the real-world venue 102. The audio demixing tool 152 and / or the audio remixing tool 154, described in further detail below, can represent one or more software tools that can be executed by one or more electrical, mechanical, and / or electromechanical devices, as will be apparent to those skilled in the art without departing from the spirit and scope of the present disclosure. Those skilled in the art will recognize that embodiments of the disclosure described herein can be implemented in hardware, firmware, software, or any combination thereof without departing from the present disclosure. Furthermore, those skilled in the art will recognize that firmware, software, routines, instructions, or the like, may be described herein as performing certain actions. However, it should be understood that such description is for convenience only and that such actions actually result from one or more electrical, mechanical, and / or electromechanical devices executing the firmware, software, routines, instructions, or the like. Alternatively, or in addition, those skilled in the art will recognize that embodiments of the disclosure described herein may also be implemented as instructions stored on a machine-readable medium, which may be read and executed by one or more processors without departing from the present disclosure. A machine-readable medium may include any mechanism for storage in a form readable by a machine, such as, by way of example, a computing device.For example, a machine-readable medium may include read-only memory (ROM), random-access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, and the like.

[0022] 1, the audio playback server 104 can execute an audio demixing tool 152 to seamlessly decompose the composite audio program 150 into audio sounds 156.1-156.n. In some embodiments, the audio sounds 156.1-156.n can include one or more audio channels, for example, two (2) audio channels, i.e., a stereophonic (stereo) audio channel that may include a left monophonic (mono) audio channel and a right audio channel feed, or five (5) audio channels, i.e., a 5.1 surround sound audio channel that may include, among others, a left monophonic (mono) audio channel, a center monophonic audio channel, a right monophonic audio channel, a left monophonic audio channel, and / or a right monophonic audio channel. In some embodiments, the audio sounds 156.1-156.n can include sounds generated by audio sources such as electronic, mechanical, and / or electromechanical devices, as would be apparent to those skilled in the art without departing from the spirit and scope of this disclosure. For example, these electronic, mechanical, and / or electromechanical devices may include, by way of example, simple instruments such as a snare drum, and / or a more complex collection of instruments such as a standard drum kit having a snare drum, a bass drum, one or more tom-toms, one or more cymbals, and / or one or more hi-hat cymbals. In some embodiments, the simple instruments may include percussion, wind, string, and / or electronic instruments, to name a few. Also, in some embodiments, the collection of instruments may include instruments from the same and / or different classes of instruments, such as percussion, wind, string, and / or electronic instruments, to name a few. Alternatively, or in addition, audio sounds 156.1-156.n may include natural audio sounds produced by non-human and / or human creatures, such as musical audio sounds produced using the human voice, often referred to as vocals.These natural audio sounds may also include natural, non-living sources, such as water and / or lightning, to name a few. In some embodiments, the audio demixing tool 152 may analyze the audio sounds 156.1-156.n to identify one or more audio sources, such as electronic, mechanical, and / or electromechanical devices, one or more non-human living beings, and / or one or more human living beings, among others, that produced these audio sounds.

[0023] As described herein, the audio demixing tool 152 can decompose the composite audio program 150 into audio voices 156.1-156.n. In some embodiments, the audio demixing tool 152 can analyze the composite audio program 150 and identify the audio voices 156.1-156.n present within the composite audio program 150. As part of this identification, the audio demixing tool 152 can iteratively search the composite audio program 150 for the audio voices 156.1-156.n. For example, the audio demixing tool 152 can search the composite audio program 150 for the audio voice produced by the snare drum, the audio voice produced by the bass guitar, etc. In some embodiments, the audio demixing tool 152 can iteratively search the composite audio program 150 starting with the most dominant or dominant audio voice among the audio voices 156.1-156.n. After identifying the audio voices 156.1-156.n, the audio demixing tool 152 may decompose the composite audio program 150 into the audio voices 156.1-156.n. As part of this decomposition, the audio demixing tool 152 may iteratively isolate the audio voices 156.1-156.n from the composite audio program 150 to provide corresponding audio voices from the audio voices 156.1-156.n. After this decomposition, the audio demixing tool 152 may store the corresponding audio voices for retrieval by the audio remixing tool 154 as described herein.

[0024] 1 , the audio remixing tool 154 can intelligently construct an audio presentation from the audio voices 156.1-156.n to play the composite audio program 150 within the real-world venue 102. In some embodiments, the audio remixing tool 154 can construct the audio presentation in advance, e.g., offline prior to the event, and / or during the event, e.g., in real time or near real time, concurrently with the event. In some embodiments, the audio presentation, when played back by the audio playback server 104, can configure the real-world venue 102 to play the audio voices 156.1-156.n on real-world loudspeakers within the real-world venue 102. In the exemplary embodiment illustrated in FIG. 1 , the audio remixing tool 154 can assign the audio voices 156.1-156.n to real-world loudspeakers within the real-world venue 102 and construct the audio presentation for playing the composite audio program 150 within the real-world venue 102. In some embodiments, the audio presentation may represent a static audio presentation whereby the assignment of audio voices 156.1-156.n to real-world loudspeakers in real-world venue 102 remains fixed or static throughout the composite audio program 150, and / or a dynamic audio presentation whereby the assignment of audio voices 156.1-156.n to real-world loudspeakers in real-world venue 102 dynamically moves, changes, or switches during the composite audio program. In some embodiments, audio remixing tool 154 may consider one or more characteristics, parameters, and / or attributes of audio voices 156.1-156.n when constructing the audio presentation.In some embodiments, these characteristics, parameters, and / or attributes can be utilized by audio remixing tool 154 to provide a more realistic, aesthetically appealing reproduction of composite audio program 150 while mitigating the effects of comb filtering and / or audio latency as described herein within real-world venue 102. Thus, mitigating the effects of comb filtering and / or audio latency as described herein within real-world venue 102 can be important to ensure a seamless and immersive experience.

[0025] After assigning the audio voices 156.1-156.n, the audio remixing tool 154 may identify audio control signals 158.1-158.i that configure the audio equipment of the real-world venue 102 to play the audio voices 156.1-156.n through real-world loudspeakers in the real-world venue 102. In some embodiments, the real-world venue 102 may include audio equipment such as amplifiers, crossovers, equalizers, and / or mixers to route and / or signal condition the audio voices 156.1-156.n for playback in the real-world venue 102. In these embodiments, the audio remixing tool 154 may generate audio control signals 158.1-158.i to configure the audio equipment in the real-world venue 102 to play the audio voices 156.1-156.n on the real-world loudspeakers in the real-world venue 102. In these embodiments, the audio control signals 158.1-158.i may cause audio equipment in the real-world venue 102 to route and / or signal condition the audio sounds 156.1-156.n for playback through real-world loudspeakers in the real-world venue 102 as specified in the audio presentation.

[0026] 1 , the real-world venue 102 may represent a three-dimensional structure, e.g., a hemispherical structure, also referred to as a hemispherical dome. In some embodiments, the real-world venue 102 may include one or more visual displays, often referred to as a three-dimensional media surface, that are spread across the interior of the real-world venue 102. In these embodiments, the one or more visual displays may include a series of rows and columns of image elements, also referred to as pixels, in three dimensions that form the three-dimensional media surface to project an image or a series of images, often referred to as video, that may be associated with an event, for example, onto the three-dimensional media surface. In these embodiments, the pixels may be implemented using one or more light-emitting diode (LED) displays, one or more organic light-emitting diode (OLED) displays, and / or one or more quantum dot (QD) displays, to name a few. For example, the three-dimensional media surface may include a 19,000 by 13,500 LED visual display that covers the perimeter of the interior of the real-world venue 102 and forms approximately 160,000 square feet of visual display.

[0027] 1, the real-world venue 102 can play audio sounds 156.1-156.n within the real-world venue 102 as defined in the audio presentation and can play a composite audio program 150 within the real-world venue 102. As shown in FIG. 1, the real-world venue 102 can include real-world loudspeakers 106.1-106.i for playing the audio sounds 156.1-156.n within the real-world venue 102. Alternatively, or in addition, the real-world venue 102 can include audio equipment as described herein for routing the audio sounds 156.1-156.n to the real-world loudspeakers 106.1-106.i and / or for signal conditioning the audio sounds 156.1-156.n. In some embodiments, the audio remixing tool 154 can play the audio sounds 156.1-156.n through the real-world loudspeakers 106.1-106.i as defined in the audio presentation in a manner substantially similar to that described herein. In some embodiments, the real-world loudspeakers 106.1-106.i can include a proscenium array real-world loudspeaker system mounted at or near the proscenium of the real-world venue 102, one or more effects extension array real-world loudspeaker systems mounted at or near the proscenium array real-world loudspeaker system, and / or one or more ambient array real-world loudspeaker systems mounted throughout the real-world venue 102. In some embodiments, the proscenium array real-world loudspeaker system, the one or more effects-expansion array real-world loudspeaker systems, and / or the one or more ambient array real-world loudspeaker systems may include one or more real-world loudspeakers, which may include one or more super tweeters, one or more tweeters, one or more mid-range speakers, one or more woofers, one or more subwoofers, and / or one or more full-range speakers, to name a few. (An exemplary audio demixing tool for decomposing a composite audio program)

[0028] FIG. 2 diagrammatically illustrates the operation of an exemplary audio demixing tool within an exemplary audio system for decomposing a composite audio program, according to some exemplary embodiments of the present disclosure. In the exemplary embodiment illustrated in FIG. 2, audio demixing tool 200 can access a composite audio program, such as composite audio program 150. As described herein, audio demixing tool 200 can seamlessly decompose the composite audio program into multiple audio voices that can be collectively played back by a real-world venue. In some embodiments, audio demixing tool 200 can represent one or more software tools that can be executed by one or more electrical, mechanical, and / or electromechanical devices, such as audio playback server 104, to seamlessly decompose the composite audio program into multiple audio voices, as will be apparent to those skilled in the art without departing from the spirit and scope of the present disclosure. Audio demixing tool 200 as described herein can represent an exemplary embodiment of audio demixing tool 152.

[0029] 2, the audio demixing tool 200 can decompose the composite audio program 202 into audio voices 208.1-208.m through a process referred to as music source separation. In some embodiments, the audio voices 208.1-208.m can include one or more audio channels, for example, two (2) audio channels, i.e., stereophonic (stereo) audio channels, which may include a left mono audio channel and a right audio channel feed, or five (5) audio channels, i.e., 5.1 surround audio channels, which may include a left mono audio channel, a center mono audio channel, a right mono audio channel, a left mono audio channel, and / or a right mono audio channel, among others. As part of this music source separation, the audio demixing tool 200 can analyze the composite audio program 202 and identify, e.g., matching audio voices 208.1-208.m from the composite audio program 202 that correspond to sample audio voices from among the sample audio voices 204.1-204.r in the electronic library of audio voices 206. The sample audio voices 204.1-204.r can include audio voices produced by one or more audio sources, such as one or more electronic, mechanical, and / or electromechanical devices, one or more non-human living beings, and / or one or more human living beings, among others, as described herein. For example, sample audio sounds 204.1-204.r may include, among others, audio sample 204.1 generated by a microphone, audio sample 204.2 generated by an acoustic guitar, an electric guitar, and a bass guitar, audio sample 204.a generated by a drum kit, audio sample 204.a+1 generated by a bass drum from the drum kit, and audio sample 204.r generated by a snare drum from the drum kit.

[0030] In some embodiments, the audio demixing tool 200 can iteratively search the composite audio program 202 for the presence of sample audio sounds 204.1-204.r. For example, the audio demixing tool 200 can iteratively search the composite audio program 202 for the presence of, among other things, an audio sample 204.1 generated by a microphone, an audio sample 204.2 generated by an acoustic guitar, an electric guitar, and a bass guitar, an audio sample 204.a generated by a drum kit, an audio sample 204.a+1 generated by a bass drum from the drum kit, and an audio sample 204.r generated by a snare drum from the drum kit. In the exemplary embodiment illustrated in FIG. 2, the audio demixing tool 200 can execute a pattern recognition algorithm to iteratively search the composite audio program 202 for sample audio sounds 204.1-204.r. In some embodiments, the pattern recognition algorithm may include a template matching algorithm that matches the sample audio sounds 204.1-204.r to the composite audio program 202, and / or a structural / syntactic matching algorithm or a statistical matching algorithm that involves semi-supervised and supervised machine learning of the sample audio sounds 204.1-204.r and their subsequent matching to the composite audio program 202, respectively.

[0031] As part of the music source separation, the audio demixing tool 200 may decompose the composite audio program 202 into audio voices 208.1-208.m. After identifying the sample audio voices 204.1-204.r, the audio demixing tool 200 may iteratively isolate each of the sample audio voices 204.1-204.r present in the composite audio program 202 from the composite audio program 202 to provide audio voices 208.1-208.m. In some embodiments, the audio demixing tool 200 may iteratively subtract each of the sample audio voices 204.1-204.r present in the composite audio program 202 from the composite audio program 202 to isolate each of the sample audio voices 204.1-204.r present in the composite audio program 202. From the above example, audio demixing tool 200 can identify, among other things, that audio sample 204.1 generated by a microphone, audio sample 204.2 generated by an acoustic guitar, an electric guitar, and a bass guitar, and audio sample 204.a generated by a drum kit are present in composite audio program 202. In this example, audio demixing tool 200 can, among other things, isolate the audio sound generated by the microphone from composite audio program 202 and provide a first audio sound from among audio sounds 208.1-208.b, isolate the audio sounds generated by the acoustic guitar, the electric guitar, and the bass guitar and provide a second audio sound from among audio sounds 208.1-208.b, and / or isolate the audio sound generated by the drum kit from composite audio program 202 and provide a b-th audio sound from among audio sounds 208.1-208.b.

[0032] In some embodiments, the audio demixing tool 200 can hierarchically decompose the composite audio program 202 into audio voices 208.1-208.m. In these embodiments, the audio demixing tool 200 can analyze the composite audio program 202 and identify sample audio voices from among the sample audio voices 204.1-204.r present in the composite audio program 202 that correspond to a collection of instruments, such as, for example, audio sample 204.a produced by a drum kit. After identifying these sample audio voices, the audio demixing tool 200 can iteratively isolate each of these sample audio voices present in the composite audio program 202 to provide an audio voice from among the audio voices 208.1-208.m that corresponds to the collection of instruments, such as, for example, audio voice 208.b. The audio demixing tool 200 may then again decompose audio speech 208.b into audio speech 208.c-208.m to provide audio speech 208.1-208.m. The audio demixing tool 200 may analyze audio speech 208.b and identify sample audio speeches from among the sample audio speeches 204.1-204.r present within audio speech 208.b. After identifying the sample audio speeches 204.1-204.r present within audio speech 208.b, the audio demixing tool 200 may iteratively isolate each of the sample audio speeches 204.1-204.r present within audio speech 208.b from audio speech 208.b to provide audio speeches 208.c-208.m from among the audio speeches 208.1-208.m.

[0033] In some embodiments, audio demixing tool 200 can analyze audio sounds 208.1-208.m to identify one or more audio sources, such as, among others, one or more electronic, mechanical, and / or electromechanical devices, one or more non-human organisms, and / or one or more human organisms, as described herein, that produced these audio sounds. In these embodiments, audio demixing tool 200 can analyze one or more of audio sounds 208.1-208.m to identify, for example, whether a simple musical instrument, such as a snare drum, produced these audio sounds. Alternatively, or in addition, audio demixing tool 200 can analyze one or more of audio sounds 208.1-208.m to identify, for example, whether a more complex collection of musical instruments, such as a standard drum kit having a snare drum, a bass drum, one or more tom-toms, one or more cymbals, and / or one or more hi-hat cymbals, produced these audio sounds. From the above example, the audio demixing tool 200 can analyze a first audio sound from among the audio sounds 208.1-208.m and identify that the audio sound was generated by a microphone; can analyze a second audio sound from among the audio sounds 208.1-208.m and identify that the audio sound was generated by an acoustic guitar, an electric guitar, and a bass guitar; and can analyze a b-th audio sound from among the audio sounds 208.1-208.m and identify that the audio sound was generated by a drum kit. An exemplary audio demixing tool for analyzing audio speech in a composite audio program.

[0034] FIG. 3 diagrammatically illustrates the operation of an exemplary audio demixing tool for analyzing a composite audio program, according to some exemplary embodiments of the present disclosure. In the exemplary embodiment illustrated in FIG. 3, audio demixing tool 300 can access multiple audio voices of a composite audio program, such as composite audio program 150. As described herein, audio demixing tool 300 can analyze the multiple audio voices and identify one or more characteristics, parameters, and / or attributes of these audio voices. As described herein, the one or more characteristics, parameters, and / or attributes can be used by an audio remixing tool, such as audio remixing tool 154, to construct an audio presentation for playing the composite audio program in a real-world venue. In some embodiments, the audio remixing tool can utilize the one or more characteristics, parameters, and / or attributes to provide a more realistic, aesthetically appealing playback of the composite audio program, while mitigating the effects of comb filtering and / or audio latency as described herein in the real-world venue. In these embodiments, mitigating the effects of comb filtering and / or audio latency as described herein within real-world venues may be important to ensure a seamless and immersive experience.

[0035] 3, the audio demixing tool 300 can access audio sounds 302.1-302.m decomposed from a composite audio program in a manner substantially similar to that described herein. In some embodiments, the audio sounds 302.1-302.m can represent sounds produced by one or more of electronic, mechanical, and / or electromechanical devices, one or more non-human organisms, and / or one or more human organisms, among others, as described herein. After accessing the audio sounds 302.1-302.m, the audio demixing tool 300 can analyze the audio sounds 302.1-302.m and identify one or more characteristics, parameters, and / or attributes 306.1-306.m corresponding to these audio sounds. 3, audio demixing tool 300 may include an audio speech analysis software engine 304 to process audio speech 302.1-302.m and identify one or more characteristics, parameters, and / or attributes 306.1-306.m. For example, audio speech analysis software engine 304 may process audio speech 302.1 and identify one or more characteristics, parameters, and / or attributes 306.1 corresponding to audio speech 302.1, and process audio speech 302.2 and identify one or more characteristics, parameters, and / or attributes 306.2 corresponding to audio speech 302.2, among other things. In some embodiments, one or more characteristics, parameters, and / or attributes 306.1-306.m may include, among other things, pitch, volume, timbre, frequency, amplitude, wavelength, and / or rate of audio speech 302.1-302.m. In these embodiments, the audio voice analysis software engine 304 may compare pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity between the audio voices 302.1-302.m and identify one or more characteristics, parameters, and / or attributes 306.1-306.m.In some embodiments, the audio speech analysis software engine 304 can store one or more characteristics, parameters, and / or attributes 306.1-306.m as an organized collection of data, often referred to as a database. The database may include one or more data tables having data values ​​such as alphanumeric strings, integers, decimals, floating-point numbers, dates, times, binary values, Boolean values, and / or enumerations, to name a few. The database can be a column-oriented database, a relational database, a keystore database, a graph database, and / or a document store, to name a few. In these embodiments, an audio remixing tool can access the database, retrieve one or more characteristics, parameters, and / or attributes 306.1-306.m, and construct an audio presentation for playing the composite audio program in a real-world venue as described herein.

[0036] In some embodiments, one or more characteristics, parameters, and / or attributes 306.1-306.m may indicate spatial positioning between one or more audio sources that produced audio sounds 302.1-302.m, audio transients of audio sounds 302.1-302.m, spatial movement of one or more audio sources that produced audio sounds 302.1-302.m, timing relationships between audio sounds 302.1-302.m, audio effects within audio sounds 302.1-302.m, and / or any other suitable characteristics, parameters, and / or attributes of or between audio sounds 302.1-302.m that would be recognized by those skilled in the art without departing from the spirit and scope of the present disclosure.

[0037] The spatial positioning between one or more audio sources that produced audio sounds 302.1-302.m indicates the relative positioning between the two or more audio sources that produced two or more of audio sounds 302.1-302.m with respect to each other. In these embodiments, the audio sound analysis software engine 304 may implement a three-dimensional audio sound localization technique between two or more of audio sounds 302.1-302.m to estimate the spatial positioning between the two or more audio sources that produced these audio sounds. In these embodiments, the audio sound analysis software engine 304 may estimate the interaural time difference (ITD) between two or more of audio sounds 302.1-302.m and / or the interaural intensity difference (IID) between two or more of audio sounds 302.1-302.m to estimate the relative positioning between the two or more audio sources with respect to each other.

[0038] The audio transients of the audio sounds 302.1-302.m represent the envelopes of the audio sounds 302.1-302.m over time. In some embodiments, these envelopes may represent the attack transients, decay transients, sustain transients, and / or release transients of the audio sounds 302.1-302.m. In these embodiments, the attack transients represent a first duration for the audio sounds 302.1-302.m to reach their maximum amplitude, the decay transients represent a second duration for the audio sounds 302.1-302.m to decrease from their maximum amplitude to their steady-state amplitude, the sustain transients represent a third duration for the audio sounds 302.1-302.m to be at their steady-state amplitude, and the release transients represent a fourth duration for the audio sounds 302.1-302.m to decrease from their steady-state amplitude to their minimum amplitude.

[0039] The spatial movement of one or more audio sources that generated the audio sounds 302.1-302.m indicates the relative movement between the two or more audio sources that generated two or more of the audio sounds 302.1-302.m relative to one another. In the exemplary embodiment illustrated in FIG. 3, the two or more audio sources may change location or move around during the composite audio program. In some embodiments, the audio sound analysis software engine 304 may analyze the audio sounds 302.1-302.m to identify whether the two or more audio sources are moving. For example, the audio demixing tool 300 may compare the amplitude and / or phase of the two or more of the audio sounds 302.1-302.m corresponding to the two or more audio sources to one another to estimate whether the two or more audio sources are moving relative to one another.

[0040] The timing relationships between the audio sounds 302.1-302.m indicate the relative timing relationship between the two or more audio sources that generated the two or more of the audio sounds 302.1-302.m relative to each other. In some embodiments, the timing relationships may be referred to as beat, bar, and / or tick relationships between the two or more audio sources. These relative timing relationships may include time signatures such as simple, compound, beat, common-time, complex, mixed, additive, irrational, and / or the like. In some embodiments, the audio analysis software engine 304 may execute a timing relationship algorithm to compare two or more of the audio sounds 302.1-302.m to each other and identify the relative timing relationship between the two or more audio sources. As part of the timing relationship algorithm, the audio analysis software engine 304 may classify the two or more audio sources. In some embodiments, the audio voice analysis software engine 304 can classify each of the two or more audio sources according to a general type, such as percussion, wind, string, and / or electronic. The audio voice analysis software engine 304 can then identify relative timing relationships between the two or more audio sources according to their general type. For example, a percussion instrument from the two or more audio sources may have the same timing relationship as another percussion instrument from the two or more audio sources, while a percussion instrument may have a different timing relationship from a wind instrument from the two or more audio sources. Alternatively, or in addition, the audio voice analysis software engine 304 can classify each of the two or more audio sources according to a specific type, such as snare drum, bass drum, tom-tom, or cymbal. The audio voice analysis software engine 304 can then identify relative timing relationships between the two or more audio sources according to their specific type.For example, a snare drum from among two or more audio sources may have the same timing relationship to another snare drum from among two or more audio sources, while a snare drum may have a different timing relationship to a cello from among two or more audio sources.

[0041] The audio effects in audio voices 302.1-302.m represent specific audio effects that may be applied by one or more audio sources that generated audio voices 302.1-302.m to modify the audio voices 302.1-302.m. In the exemplary embodiment illustrated in FIG. 3, the audio voice analysis software engine 304 can analyze the audio voices 302.1-302.m and identify audio effects within the audio voices 302.1-302.m. In some embodiments, the one or more audio effects can include one or more echoes, flangers, phasers, choruses, equalization, filtering, overdrives, pitch shifts, time stretches, resonators, voice effects, synthesizers, modulations, compressions, and / or the like, to name a few. In some embodiments, the audio voice analysis software engine 304 can isolate the audio voice being generated by the audio source, referred to as the parent audio voice, from the audio effects within the audio voices 302.1-302.m, referred to as the child audio voices. Example Operation of an Example Audio Demixing Tool

[0042] FIG. 4 illustrates a flowchart of an exemplary audio demixing tool according to some exemplary embodiments of the present disclosure. This disclosure is not limited to this operational description. Rather, other operational control flows will be apparent to those skilled in the art and are within the scope and spirit of the present disclosure. The following discussion describes an operational control flow 400 for seamlessly decomposing a composite audio program, such as composite audio program 150, and / or identifying characteristics, parameters, and / or attributes corresponding to audio voices of the composite audio program. The operational control flow 400 can be implemented by, for example, audio demixing tool 152. In some embodiments, the operational control flow 400 can be executed by one or more computing devices, such as audio playback server 104.

[0043] At operation 402, operation control flow 400 decomposes the composite audio program in a manner substantially similar to that described herein to provide audio sounds such as audio sounds 156.1-156.n, audio sounds 208.1-208.m, and / or audio sounds 302.1-302.m.

[0044] At operation 404, operation control flow 400 may analyze one or more of the audio sounds from operation 402 and identify one or more characteristics, parameters, and / or attributes of those audio sounds in a manner substantially similar to that described herein. Also, as described herein, these characteristics, parameters, and / or attributes may be utilized by an audio remixing tool, such as audio remixing tool 154, to construct an audio presentation for playing the composite audio program in a real-world venue. Example Operation of an Example Audio Remixing Tool

[0045] Before describing an example audio remixing tool that may be implemented within the example real-world venue described herein, audio voice ensembles will be generally described. As described herein, a composite audio program may include multiple audio voices generated by multiple audio sources, such as percussion instruments, wind instruments, string instruments, and / or electronic instruments, to name a few. As described herein, these audio sources may be logically grouped together to form an audio voice ensemble. In some embodiments, audio sources from among multiple audio sources having similar characteristics, parameters, and / or attributes may be logically grouped together to form an audio voice ensemble. In these embodiments, audio sources from among multiple audio sources having similar pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity, to name a few, may be logically grouped together to form an audio voice ensemble. For example, a snare drum, a bass drum, one or more tom-toms, one or more cymbals, and / or one or more hi-hat cymbals having similar pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity can be logically grouped together to form an audio voice ensemble associated with a drum kit.

[0046] FIG. 5 graphically illustrates an exemplary ensemble bounding volume that may be generated by an exemplary audio remixing tool in an exemplary audio system, according to some exemplary embodiments of the present disclosure. In the exemplary embodiment illustrated in FIG. 5, multiple audio sources having similar characteristics, parameters, and / or attributes may be logically grouped together to form audio voice ensembles. As described herein, these audio voice ensembles may be associated with ensemble bounding volumes. These ensemble bounding volumes may define the spatial distance between audio sources in the audio voice ensemble to provide a more realistic, aesthetically appealing reproduction of a composite audio program, such as composite audio program 150, while mitigating the effects of comb filtering and / or audio latency as described herein in a real-world venue, such as real-world venue 102. Thus, mitigating the effects of comb filtering and / or audio latency as described herein in a real-world venue may be important to ensure a seamless and immersive experience. 5 describes an example ensemble bounding volume for a drum kit, which may include a snare drum, a bass drum, one or more tom-toms, one or more cymbals, and / or one or more hi-hat cymbals, to name a few. Those skilled in the art will recognize that other ensemble bounding volumes for drum kits and / or other ensemble bounding volumes for other audio sources may be implemented in a substantially similar manner to that described herein without departing from the spirit and scope of the present disclosure.

[0047] As illustrated in FIG. 5 , audio sources 500.1-500.x having similar characteristics, parameters, and / or attributes from among multiple audio sources of a composite audio program, such as, for example, a snare drum, a bass drum, one or more tom-toms, one or more cymbals, and / or one or more hi-hat cymbals of a drum kit, can be logically grouped together to form an audio voice ensemble 502. In the exemplary embodiment illustrated in FIG. 5 , audio voice ensemble 502 having audio sources 500.1-500.x can be associated with an ensemble bounding volume 504. While ensemble bounding volume 504 is illustrated in FIG. 5 as being a bounding box in three-dimensional space, this is for illustrative purposes and is not intended to be limiting. Those skilled in the art will recognize that ensemble bounding volume 504 can be any suitable three-dimensional volume in three-dimensional space, such as a bounding capsule, a bounding cylinder, a bounding ellipsoid, a bounding sphere, a bounding slab, and / or a bounding triangle, without departing from the spirit and scope of the present disclosure. Also, while ensemble bounding volume 504 is illustrated in FIG. 5 as being in three-dimensional space, those skilled in the art will recognize that ensemble bounding volume 504 may similarly be implemented as any suitable two-dimensional shape in two-dimensional space, such as a circle, triangle, quadrilateral, and / or polygon, to form an ensemble bounding area without departing from the spirit and scope of this disclosure. In some embodiments, audio sources 500.1-500.x can be characterized as being diffuse audio sources, transient audio sources, and / or any combination of diffuse and transient audio sources. In these embodiments, diffuse audio sources represent audio sources that produce sounds that occur gradually over a relatively long duration, while transient audio sources produce sounds that occur suddenly over a relatively short duration.

[0048] Generally, the ensemble bounding volume 504 defines the spatial distance between the audio sources 500.1-500.x in three-dimensional space. In the exemplary embodiment illustrated in FIG. 5, the ensemble bounding volume 504 defines the maximum spatial distance between the audio sources 500.1-500.x in three-dimensional space. In the exemplary embodiment illustrated in FIG. 5, the maximum spatial distance between the audio sources 500.1-500.x represents a spatial distance between the audio sources 500.1-500.x that has a difference between the audible times of flight that is less than or equal to the audible time-of-flight threshold. In some embodiments, the audible time-of-flight threshold is less than or equal to the first audible time-of-flight threshold. A,1 and the second audible time of flight T B,1 can be selectively chosen such that the difference between the first audible time-of-flight T is less than the time resolution of human hearing, e.g., between about twenty (20) milliseconds and about thirty-six (36) milliseconds. In these embodiments, the effects of comb filtering and / or audio latency are typically less than the first audible time-of-flight T A,1 and the second audible time of flight T B,1 is not noticeable when the difference between is less than the time resolution of human hearing as described herein.

[0049] As illustrated in FIG. 5, the audio demixing tool determines a first three-dimensional coordinate (X A ,Y A ,Z A ), and a second three-dimensional coordinate (X B ,Y B ,Z B ) distance D between A,B The audio demixing tool can then identify that the first audio sound generated by audio source 500.1 is located at a first three-dimensional coordinate (X A ,Y A ,Z A ) to the first 3D coordinate (x A ,y A ,z A ) the first audible time of flight TA,1 and a second audio sound generated by audio source 500.x is located at a second three-dimensional coordinate (X B ,Y B ,Z B ) to the first 3D coordinate (x A ,y A ,z A ) the second audible time of flight T B,1 In some embodiments, the audio demixing tool may estimate a first audible time of flight T A,1 can be roughly estimated as follows: [ka] Second audible time of flight T B,1 can be roughly estimated as follows: [ka] T1 and T2 are the first audible time-of-flight T A,1 and the second audible time of flight T B,1 D1 and D2 represent the first three-dimensional coordinate (X A ,Y A ,Z A ) and the first three-dimensional coordinate (x A ,y A ,z A ) and the second three-dimensional coordinate (X B ,Y B ,Z B ) and the first three-dimensional coordinate (x A ,y A ,z A ) and v sound represents the speed of sound. Typically, the speed of sound in air is approximately three hundred and forty-three (343) meters per second at twenty (20) degrees Celsius, which can vary depending on temperature. The audio demixing tool then calculates the first audible time of flight, T A,1 and the second audible time of flight T B,1 In some embodiments, the audio demixing tool may estimate the difference between the first audible time of flight TA,1 and the second audible time of flight T B,1 The difference between the first audible time of flight T can be compared to an audible time of flight threshold. In these embodiments, the audio demixing tool may A,1 and the second audible time of flight T B,1 is less than or equal to the audible time-of-flight threshold, audio source 500.1 and audio source 500.x are spaced apart by a distance D. A,B In these embodiments, the distance D A,B is the first audible time of flight T A,1 and the second audible time of flight T B,1 is less than or equal to the audible time-of-flight threshold. Otherwise, in some embodiments, the audio demixing tool may determine that the first audible time-of-flight T A,1 and the second audible time of flight T B,1 When the difference between the audio sources 500.1 and 500.x exceeds an audible time-of-flight threshold, the audio source 500.1 and the audio source 500.x are moved to a distance D A,B In these embodiments, the distance D A,B is the first audible time of flight T A,1 and the second audible time of flight T B,1 is less than or equal to the audible time-of-flight threshold, audio sources 500.1 and 500.x may be characterized as being outside ensemble bounding volume 504. In some embodiments, the audio demixing tool may empirically simulate audio sources 500.1 and 500.x at different three-dimensional coordinates to determine the faces, vertices, and / or surfaces of ensemble bounding volume 504 in three-dimensional space. In these embodiments, the audio demixing tool may execute a computational algorithm, e.g., a Monte Carlo algorithm, to empirically simulate audio sources 500.1 and 500.x at different three-dimensional coordinates.

[0050] The audio demixing tool may also demultiplex a first audio sound generated by audio source 500.1 at a first three-dimensional coordinate (X 1 , X 2 , X 3 ) in three-dimensional space in a manner substantially similar to that described herein. A ,Y A ,Z A ) to the second 3D coordinate (x B ,y B ,z B ) the first audible time of flight T A,2 and a second audio sound generated by audio source 500.x is located at a second three-dimensional coordinate (X B ,Y B ,Z B ) to the second 3D coordinate (x B y B ,z B ) the second audible time of flight T B,2 In some embodiments, the first audible time of flight T A,2 and the second audible time of flight T B,2 The difference between the first audible time of flight T A,1 and the second audible time of flight T B,1 In these embodiments, the first three-dimensional coordinate (x A ,y A ,z A ) and a second audio sound generated by audio source 500.x at a second three-dimensional coordinate (x B y B ,z B ) in a real-world venue. Thus, the first audio sound produced by audio source 500.1 and the second audio sound produced by audio source 500.x should be substantially similar to each other in most, if not all, locations within the real-world venue. Exemplary Virtual Venues That May Be Accessed by Exemplary Audio Demixing Tools

[0051] FIG. 6 diagrammatically illustrates an exemplary virtual venue that may be accessed by an exemplary audio remixing tool according to some exemplary embodiments of the present disclosure. As described herein, an audio remixing tool, such as audio remixing tool 154, can intelligently construct an audio presentation and play a composite audio program, such as composite audio program 150, on real-world loudspeakers in a real-world venue, such as a real-world venue. As described herein, the audio remixing tool can access a virtual representation of the real-world venue in three-dimensional space, also referred to as virtual venue 600, that virtually identifies the three-dimensional coordinates of the real-world loudspeakers in the three-dimensional space. While virtual venue 600 is illustrated in FIG. 6 as being in three dimensions within three-dimensional space, those skilled in the art will recognize that virtual venue 600 may similarly be in two dimensions within two-dimensional space without departing from the spirit and scope of the present disclosure. Additionally, as described herein, an audio remixing tool can utilize the virtual venue 600 to assign audio voices of a composite audio program, such as audio voices 156.1-156.n of the composite audio program 150, to real-world loudspeakers in the real-world venue and construct an audio presentation for playing the composite audio program in the real-world venue.

[0052] As illustrated in FIG. 6 , virtual venue 600 includes virtual loudspeakers 602.1-602.k that are installed within the three-dimensional space of virtual venue 600. However, the configuration and arrangement of virtual loudspeakers 602.1-602.k within the three-dimensional space of virtual venue 600 as illustrated in FIG. 6 is for illustrative purposes and is not intended to be limiting. Those skilled in the art will recognize that real-world loudspeakers 602.1-602.k may be configured and arranged differently within the three-dimensional space of virtual venue 600 without departing from the spirit and scope of the present disclosure. In some embodiments, virtual loudspeakers 602.1-602.k are respectively arranged at three-dimensional coordinates (x1, y1, z1)-(x k ,y k ,z k) In some embodiments, the virtual loudspeakers 602.1-602.k may represent virtual representations of real-world loudspeakers in a real-world venue, such as real-world venue 102, virtual representations of virtual loudspeakers in a real-world venue, and / or any combination of real-world or virtual loudspeakers in a real-world venue. In these embodiments, the virtual representations of real-world loudspeakers may include a proscenium virtual loudspeaker system 604 mounted at or near the proscenium of the virtual venue 600, effects-enhanced virtual array real-world loudspeaker systems 608.1-608.1 mounted at or near the proscenium virtual loudspeaker system 604, and / or environmental virtual array real-world loudspeaker systems 610.1-610.m mounted throughout the virtual venue 600. In these embodiments, proscenium virtual loudspeaker system 604 can include virtual loudspeakers 606.1-606.z. In some embodiments, virtual loudspeakers 606.1-606.z, effects-enhanced virtual array real-world loudspeaker systems 608.1-608.l, and / or environmental virtual array real-world loudspeaker systems 610.1-610.m can include one or more virtual super tweeters, one or more virtual tweeters, one or more virtual midrange speakers, one or more virtual woofers, one or more virtual subwoofers, and / or one or more virtual full-range speakers, to name a few. (Example static audio presentation that can be constructed by an example audio demixing tool)

[0053] 7A-7F, described in further detail below, represent an audio presentation in which the assignment of audio voices of a composite audio program, such as audio voices 156.1-156.n of composite audio program 150, to real-world loudspeakers in a real-world venue, such as real-world venue 102, remains fixed or static throughout the composite audio program, whereas the exemplary dynamic audio presentation, described herein in FIGS. 8A-8B, represent an audio presentation in which the assignment of audio voices of a composite audio program to real-world loudspeakers in a real-world venue dynamically moves, changes, or switches during the composite audio program.

[0054] 7A-7F diagrammatically illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation for playing a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. In the exemplary embodiment illustrated in FIG. 7A, an audio remixing tool, such as audio remixing tool 154, can intelligently construct static audio presentation 700 to play the composite audio program on real-world loudspeakers in a real-world venue, such as real-world venue 102. As described herein, the audio remixing tool can utilize virtual venue 600 to assign audio voices of the composite audio program, such as audio voices 156.1-156.n of composite audio program 150, to real-world loudspeakers in the real-world venue to construct static audio presentation 700. Those skilled in the art will recognize that static audio presentation 700 as illustrated in FIG. 7A is for illustrative purposes and is not intended to be limiting. Those skilled in the art will recognize that audio remixing tools may intelligently construct other static audio presentations and play other composite audio programs on real-world loudspeakers in real-world venues without departing from the spirit and scope of the present disclosure.

[0055] 7A-7F illustrate exemplary operations that may be utilized by an audio remixing tool to construct static audio presentation 700 for playback of a composite audio program on real-world loudspeakers in a real-world venue. Those skilled in the art will recognize that these operations may be performed independently or in any combination to construct other static audio presentations for playback of other composite audio programs on real-world loudspeakers in real-world venues without departing from the spirit and scope of this disclosure. In the exemplary embodiment illustrated in FIG. 7A, the audio remix can access an electronic library of audio voices 702, which includes audio voices 704.1-704.m that are present in the composite audio program. In some embodiments, an audio demixing tool, such as audio demixing tool 152, can seamlessly decompose the composite audio program and identify audio voices 704.1-704.m. 7A, audio sounds 704.1-704.m may include audio sounds generated by one or more audio sources, such as one or more electronic, mechanical, and / or electromechanical devices, one or more non-human living beings, and / or one or more human living beings, among others, as described herein. For example, audio sounds 704.1-704.m may include audio sounds generated by a microphone 704.1, audio sounds generated by an acoustic guitar, an electric guitar, and a bass guitar 704.2, audio sounds generated by a bass drum 704.3, and audio sounds generated by a snare drum 704.m, among others.

[0056] 7A , an audio remixing tool can assign audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k in a virtual venue 600 to construct a static audio presentation 700 for playing a composite audio program on real-world loudspeakers in a real-world venue. As described herein, the virtual venue 600 can include a proscenium virtual loudspeaker system 604, effects-enhanced virtual array real-world loudspeaker systems 608.1-608.l, and / or environmental virtual array real-world loudspeaker systems 610.1-610.m, with virtual loudspeakers 606.1-606.z. As illustrated in FIG. 7A, the audio remixing tool can assign audio voices 704.1 generated by a microphone to virtual loudspeaker 606.4 from within the proscenium virtual loudspeaker system 604, audio voices 704.2 generated by an acoustic guitar, an electric guitar, and a bass guitar to virtual loudspeaker 606.7 from within the proscenium virtual loudspeaker system 604, audio voices 704.3 generated by a bass drum to virtual loudspeaker 606.2 from within the proscenium virtual loudspeaker system 604, and / or audio voices 704.m generated by a snare drum to virtual loudspeaker 606.1 from within the proscenium virtual loudspeaker system 604 to construct a static audio presentation 700. Those skilled in the art will recognize that the assignment of audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k in virtual venue 600 as illustrated in Figure 7A is for illustrative purposes and is not intended to be limiting. Those skilled in the art will recognize that audio remixing tools may assign audio voices 704.1-704.m to other virtual loudspeakers 602.1-602.k in virtual venue 600 to construct static audio presentation 700 without departing from the spirit and scope of the present disclosure.

[0057] The following discussion of Figures 7B-7F describes example operations to be performed by an audio remixing tool to assign audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k in virtual venue 600 as illustrated in Figure 7A. In some embodiments, the example operations described in more detail below in Figures 7B-7D and / or the example operations described in more detail below in Figures 8A-8B can form or be included within a software remixing toolkit that can be utilized by an audio remixing tool to construct example audio presentations described herein, e.g., static audio presentation 700. In these embodiments, the example operations within the software remixing toolkit can be performed independently or in any combination to construct these audio presentations.

[0058] 7B graphically illustrates a simple instrument assignment operation 710 that may be performed by an audio remixing tool to assign audio sounds produced by simple instruments, such as percussion, wind, string, and / or electronic instruments, to virtual loudspeakers 602.1-602.k in virtual venue 600. Generally, audio sounds produced by simple instruments such as those illustrated in FIG. 7B represent audio sounds having different characteristics, parameters, and / or attributes than other audio sounds from among audio sounds 704.1-704.m. In some embodiments, audio sounds produced by simple instruments have dissimilar pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity, to name a few, from other audio sounds from among audio sounds 704.1-704.m. 7B , the audio remixing tool can assign an audio sound generated by a simple instrument, such as, for example, audio sound 704.1 generated by a microphone, to virtual loudspeaker 712 from among virtual loudspeakers 606.1-606.z. In some embodiments, virtual loudspeaker 712 can be selected from among virtual loudspeakers 606.1-606.z based on a composite audio program, such as composite audio program 150, being played by virtual venue 600. In some embodiments, virtual loudspeaker 712 can be selected from among virtual loudspeakers 606.1-606.z through algorithmic best source decoding for virtual loudspeakers 606.1-606.z. In these embodiments, algorithmic best source decoding will be apparent to those skilled in the art without departing from the spirit and scope of this disclosure. Although audio sounds produced by simple musical instruments are described in FIG. 7B as being assigned to virtual loudspeaker 712, those skilled in the art will recognize that audio sounds produced by simple musical instruments may be assigned to any of virtual loudspeakers 602.1-602.k in virtual venue 600 in a manner substantially similar to that described in FIG. 7B without departing from the spirit and scope of the present disclosure.In some embodiments, the virtual loudspeaker 712 has three-dimensional coordinates (x, y) in the three-dimensional space of the virtual venue 600. v ,y v ,z v As described herein, the three-dimensional coordinates (x v ,y v ,z v ) can be characterized as providing a frame of reference or virtual coordinate system for the virtual loudspeakers 712 in the virtual venue 600.

[0059] 7C graphically illustrates an instrument collection assignment operation 720 that may be performed by an audio remixing tool to assign audio sounds produced by a collection of instruments, such as instruments from the same instrument category, such as percussion, wind, string, and / or electronic instruments, and / or from different instrument categories, to virtual loudspeakers 602.1-602.k in virtual venue 600. Generally, audio sounds produced by a collection of instruments such as those illustrated in FIG. 7C represent audio sounds having different characteristics, parameters, and / or attributes than other audio sounds from among audio sounds 704.1-704.m. In some embodiments, the audio sounds produced by the collection of instruments have pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity, to name a few, dissimilar to other audio sounds from among audio sounds 704.1-704.m. 7C , the audio remixing tool can assign audio sounds produced by a collection of instruments, such as audio sounds 704.2 produced by an acoustic guitar, an electric guitar, and a bass guitar, to virtual loudspeaker 722 from among virtual loudspeakers 606.1-606.z. In some embodiments, virtual loudspeaker 722 can be selected from among virtual loudspeakers 606.1-606.z based on a composite audio program, such as composite audio program 150, being played by virtual venue 600. In some embodiments, virtual loudspeaker 722 can be selected from among virtual loudspeakers 606.1-606.z through algorithmic best-source decoding for virtual loudspeakers 606.1-606.z. In these embodiments, algorithmic best-source decoding will be apparent to those skilled in the art without departing from the spirit and scope of this disclosure.Although the audio sounds produced by the collection of musical instruments are described in FIG. 7C as being assigned to virtual loudspeaker 722, those skilled in the art will recognize that the audio sounds produced by the collection of musical instruments may be assigned to any of virtual loudspeakers 602.1-602.k in virtual venue 600 in a manner substantially similar to that described in FIG. 7C without departing from the spirit and scope of the present disclosure. In some embodiments, virtual loudspeaker 722 is represented by a three-dimensional coordinate (x, y) within the three-dimensional space of virtual venue 600. G ,y G ,z G As described herein, the three-dimensional coordinates (x G ,y G ,z G ) can be characterized as providing a frame of reference or virtual coordinate system for the virtual loudspeakers 722 in the virtual venue 600.

[0060] 7D graphically illustrates an audio voice ensemble assignment operation 730 that may be performed by an audio remixing tool to assign audio voices produced by an audio voice ensemble having a collection of instruments, such as instruments from the same instrument class, such as percussion instruments, wind instruments, string instruments, and / or electronic instruments, and / or from different instrument classes, to virtual loudspeakers 602.1-602.k in virtual venue 600. Generally, the audio voices produced by an audio voice ensemble represent audio voices from among audio voices 704.1-704.m that have similar characteristics, parameters, and / or attributes to one another. In some embodiments, the audio voices produced by an audio voice ensemble have similar pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity to one another, to name a few. In the exemplary embodiment illustrated in Figure 7D, the audio remixing tool can assign an audio voice generated by a simple instrument from among the audio voice ensemble, such as, for example, audio voice 704.3 generated by a bass drum, to virtual loudspeaker 732 from among virtual loudspeakers 606.1-606.z, in a manner substantially similar to that described herein. Although the audio voice generated by the simple instrument is described in Figure 7D as being assigned to virtual loudspeaker 712, those skilled in the art will recognize that the audio voice generated by the simple instrument may be assigned to any of virtual loudspeakers 602.1-602.k in virtual venue 600 in a manner substantially similar to that described in Figure 7D without departing from the spirit and scope of the present disclosure. The audio remixing tool can then analyze the audio voices 704.1-704.m and identify that another audio voice produced by another simple instrument, such as, for example, audio voice 704.m produced by a snare drum, has similar characteristics, parameters, and / or attributes as the audio voice produced by the simple instrument, such as, for example, audio voice 704.3 produced by a bass drum.In some embodiments, the audio remixing tool can access an ensemble bounding volume 504 as described herein that defines the spatial distance between the simple instrument and other simple instruments within the three-dimensional space of the virtual venue 600. After accessing the ensemble bounding volume 504, the audio remixing tool can assign audio sounds produced by other simple instruments from within the audio sound ensemble, such as, for example, audio sound 704.m produced by a snare drum, to virtual loudspeakers 734 spatially positioned within the ensemble bounding volume 504 from among the virtual loudspeakers 606.1-606.z, in a manner substantially similar to that described herein. In some embodiments, virtual loudspeaker 732 and virtual loudspeaker 734 each have three-dimensional coordinates (x, z) within the three-dimensional space of the virtual venue 600. BD ,y BD ,z BD ) and three-dimensional coordinates (x SD ,y SD ,z SD As described herein, the three-dimensional coordinates (x BD ,y BD ,z BD ) and three-dimensional coordinates (x SD ,y SD ,z SD ) can be characterized as providing a frame of reference or virtual coordinate system for the virtual loudspeaker 732 and the virtual loudspeaker 734, respectively, within the virtual venue 600. In some embodiments, the audio remixing tool calculates the distance between an audio sound produced by a simple instrument, such as, for example, audio sound 704.3 produced by a bass drum, and an audio sound produced by another simple instrument, such as, for example, audio sound 704.m produced by a snare drum, i.e., the three-dimensional coordinate (x BD ,y BD ,z BD ) and three-dimensional coordinates (x SD ,y SD ,z SD) can be estimated. In these embodiments, the audio remixing tool can compare this distance to the ensemble bounding volume 504 to verify that the audio sounds produced by the simple instrument and the audio sounds produced by other simple instruments are within the area of ​​the ensemble bounding volume 504.

[0061] As described herein in FIGS. 7E and 7F, one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m may affect the allocation of the audio voices 704.1-704.m to the virtual loudspeakers 602.1-602.k within the virtual venue 600. In some embodiments, the one or more characteristics, parameters, and / or attributes may indicate spatial positioning between audio sources, such as electronic, mechanical, and / or electromechanical devices, one or more non-human living beings, and / or one or more human living beings, among others, that generated audio sounds 704.1-704.m; audio transients of audio sounds 704.1-704.m; spatial movement of audio sources that generated audio sounds 704.1-704.m; timing relationships between audio sounds 704.1-704.m; audio effects within audio sounds 704.1-704.m; and / or any other suitable characteristics, parameters, and / or attributes of or between audio sounds 704.1-704.m, as would be recognized by one of ordinary skill in the art without departing from the spirit and scope of the present disclosure.

[0062] In some embodiments, the audio remixing tool can access one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m. In these embodiments, the audio remixing tool can analyze the audio voices 704.1-704.m and identify one or more characteristics, parameters, and / or attributes of these audio voices in a manner substantially similar to that described herein. Alternatively, or in addition, the audio remixing tool can access a database as described herein and retrieve one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m. After accessing the one or more characteristics, parameters, and / or attributes, the audio remixing tool can assign the audio voices produced by the collection of musical instruments to the virtual loudspeakers 602.1-602.k in the virtual venue 600 according to the characteristics, parameters, and / or attributes of the audio voices 704.1-704.m.

[0063] 7E diagrammatically illustrates audio effect operations 740 that may be performed by an audio remixing tool to assign audio voices produced by simple instruments, such as percussion, wind, string, and / or electronic instruments, and / or collections of instruments, such as instruments from the same and / or different instrument classes, such as percussion, wind, string, and / or electronic instruments, to virtual loudspeakers 602.1-602.k in virtual venue 600. As described herein, the audio remixing tool may access one or more characteristics, parameters, and / or attributes of audio voices 704.1-704.m, such as, for example, audio effects within audio voices 704.1-704.m. Alternatively, or in addition, the audio remixing tool can analyze the audio voices 704.1-704.m and identify one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m in a manner substantially similar to that described herein. In the exemplary embodiment illustrated in FIG. 7E, the audio remixing tool can isolate audio voices produced by simple instruments and / or collections of instruments, referred to as parent audio voices, from audio effects within the audio voices 704.1-704.m, referred to as child audio voices. In some embodiments, these audio effects can include one or more echoes, flangers, phasers, choruses, equalization, filtering, overdrives, pitch shifts, time stretches, resonators, voice effects, synthesizers, modulations, compressions, and / or the like, to name a few. As shown in FIG. 7E, the audio remixing tool can isolate an audio voice from within the audio voice 704.m produced by the snare drum, referred to as the parent audio voice 742 in FIG. 7E, from the audio effects within the audio voice 704.m produced by the snare drum, referred to as the child audio voice 744 in FIG. 7E.

[0064] After isolating the parent and child audio sounds, the audio remixing tool can identify a parent-child real-world loudspeaker pairing from among the virtual loudspeakers 602.1-602.k in the virtual venue 600. In some embodiments, the parent-child real-world loudspeaker pairing can include a virtual loudspeaker 746 from among the virtual loudspeakers 606.1-606.z associated with an effect-augmented virtual array real-world loudspeaker system 748 from among the effect-augmented virtual array real-world loudspeaker systems 608.1-608.l. In some embodiments, the audio remixing tool can utilize a predetermined set of source separation rules to identify the parent-child real-world loudspeaker pairing. In these embodiments, these source separation rules can be based on transients and / or diffuse audio sounds within the parent and / or child audio sounds. For example, the parent-child real-world loudspeaker pairing can be based on diffuse audio sounds within the parent and / or child audio sounds. In some embodiments, the predetermined set of source separation rules preferably maintains, as much as possible, the angle and / or distance relationship between the parent and child audio sounds.

[0065] In the exemplary embodiment illustrated in FIG. 7E , the audio remixing tool can assign parent audio voices generated by simple instruments and / or collections of instruments, such as parent audio voice 742, to virtual loudspeaker 746 from among virtual loudspeakers 606.1-606.z, and the audio remixing tool can assign child audio voices generated by simple instruments and / or collections of instruments, such as child audio voice 744, to effect-augmented virtual array real-world loudspeaker systems 748 from among effect-augmented virtual array real-world loudspeaker systems 608.1-608.l. Although parent audio voices produced by simple instruments and / or collections of instruments are described as being assigned to virtual loudspeaker 746, and child audio voices produced by simple instruments and / or collections of instruments are described as being assigned to effect-augmented virtual array real-world loudspeaker system 748 of FIG. 7E, those skilled in the art will recognize that the parent audio voices and / or child audio voices may be assigned to any of virtual loudspeakers 602.1-602.k in virtual venue 600 in a manner substantially similar to that described in FIG. 7E without departing from the spirit and scope of the present disclosure. In some embodiments, virtual loudspeaker 746 and effect-augmented virtual array real-world loudspeaker system 748 are each assigned to three-dimensional coordinates (x, y, z) in the three-dimensional space of virtual venue 600. SDPARENT ,y SDPARENT ,z SDPARENT ) and three-dimensional coordinates (x SDCHILD ,y SDCHILD ,z SDCHILD As described herein, the three-dimensional coordinates (x SDPARENT ,y SDPARENT ,z SDPARENT ) and three-dimensional coordinates (x SDCHILD ,y SDCHILD ,z SDCHILD ) can be characterized as providing a frame of reference or virtual coordinate system for the virtual loudspeaker 746 and the effects-enhanced virtual array real-world loudspeaker system 748, respectively.

[0066] 7F diagrammatically illustrates audio transient operations 750 that may be performed by an audio remixing tool to assign audio voices produced by simple instruments, such as percussion, wind, string, and / or electronic instruments, and / or collections of instruments, such as instruments from the same and / or different instrument classes, such as percussion, wind, string, and / or electronic instruments, to virtual loudspeakers 602.1-602.k in virtual venue 600. As described herein, the audio remixing tool may have access to one or more characteristics, parameters, and / or attributes of audio voices 704.1-704.m, such as, by way of example, audio transients of audio voices 704.1-704.m. Alternatively, or in addition, the audio remixing tool may analyze the audio sounds 704.1-704.m and identify one or more characteristics, parameters, and / or attributes of the audio sounds 704.1-704.m in a manner substantially similar to that described herein.

[0067] In the exemplary embodiment illustrated in Figure 7E, the audio remixing tool can isolate the attack transient from among the audio audio 704.1-704.m and the decay transient from among the audio audio 704.1-704.m. As illustrated in Figure 7F, the audio remixing tool can isolate the attack transient from among the audio audio 704.m generated by the snare drum, referred to as attack transient audio audio 752 in Figure 7F, and the decay transient from among the audio audio 704.m generated by the snare drum, referred to as decay transient audio audio 754 in Figure 7F. In some embodiments, the attack transient audio audio 752 represents a first duration for the snare drum to reach its maximum amplitude, and the decay transient audio audio 754 represents a second duration for the snare drum to decrease from its maximum amplitude to its steady-state amplitude.

[0068] After isolating the attack and decay transients, the audio remixing tool can identify an attack-decay real-world loudspeaker pairing from among the virtual loudspeakers 602.1-602.k in the virtual venue 600. In some embodiments, the attack-decay real-world loudspeaker pairing can include a virtual loudspeaker 756 from among the virtual loudspeakers 606.1-606.z associated with an effect-enhanced virtual array real-world loudspeaker system 758 from among the effect-enhanced virtual array real-world loudspeaker systems 608.1-608.l.

[0069] In the exemplary embodiment illustrated in FIG. 7F , the audio remixing tool can assign attack transients generated by simple instruments and / or collections of instruments, such as attack transient audio voice 752, to a virtual loudspeaker 756 from among virtual loudspeakers 606.1-606.z, and the audio remixing tool can assign decay transients generated by simple instruments and / or collections of instruments, such as decay transient audio voice 754, to an effect-augmented virtual array real-world loudspeaker system 758 from among effect-augmented virtual array real-world loudspeaker systems 608.1-608.l. Although attack transients produced by simple instruments and / or collections of instruments are described as being assigned to virtual loudspeaker 756, and decay transients produced by simple instruments and / or collections of instruments are described as being assigned to effect-augmented virtual array real-world loudspeaker system 758 of FIG. 7F, those skilled in the art will recognize that attack and / or decay audio sounds may be assigned to any of virtual loudspeakers 602.1-602.k in virtual venue 600 in a manner substantially similar to that described in FIG. 7F without departing from the spirit and scope of the present disclosure. In some embodiments, virtual loudspeaker 756 and effect-augmented virtual array real-world loudspeaker system 758 are each assigned to three-dimensional coordinates (x, y, z) within the three-dimensional space of virtual venue 600. SDATTACK ,y SDATTACK ,z SDATTACK ) and three-dimensional coordinates (x SDDECAY ,y SDDECAY ,z SDDECAY As described herein, the three-dimensional coordinates (x SDATTACK ,y SDATTACK ,z SDATTACK ) and three-dimensional coordinates (x SDDECAY ,y SDDECAY ,z SDDECAY) can be characterized as providing a frame of reference or virtual coordinate system for the virtual loudspeaker 756 and the effects-enhanced virtual array real-world loudspeaker system 758, respectively.

[0070] 7A-7F , one or more characteristics, parameters, and / or attributes of the real-world venue, such as the seating arrangement within the real-world venue, the location of the performance stage within the real-world venue, and / or the location of the real-world loudspeakers within the real-world venue, may influence the allocation of audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k in the virtual venue 600. For example, the allocation of audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k in the virtual venue 600 may be based on time domain, volume level, coverage uniformity, and frequency bandwidth, adjusted by an analysis of the artist's original positioning intent. In some embodiments, the seating arrangement within the real-world venue may dictate the amount of influence on the allocation of audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k in the virtual venue 600. In some embodiments, the allocation of audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k within the virtual venue 600 can be determined by the creative artist's intent, the transient nature of the content, and the temporal relationship of the content to other content elements. Exemplary Dynamic Audio Presentations That May Be Constructed by Exemplary Audio Demixing Tools

[0071] 8A-8B diagrammatically illustrate the operation of an exemplary audio remixing tool in constructing an exemplary dynamic audio presentation for playing a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. In the exemplary embodiment illustrated in FIG. 8A, an audio remixing tool, such as audio remixing tool 154, can intelligently construct dynamic audio presentation 800 to play the composite audio program on real-world loudspeakers in a real-world venue, such as real-world venue 102, as described herein. As described herein, the audio remixing tool can utilize virtual venue 600, as described herein, to assign audio voices of the composite audio program, such as audio voices 156.1-156.n of composite audio program 150, as described herein, to real-world loudspeakers in the real-world venue to construct dynamic audio presentation 800. Those skilled in the art will recognize that dynamic audio presentation 800 as illustrated in FIG. 8A is for illustrative purposes and is not intended to be limiting. Those skilled in the art will recognize that audio remixing tools can intelligently construct other dynamic audio presentations and play other composite audio programs on real-world loudspeakers in real-world venues without departing from the spirit and scope of the present disclosure.

[0072] 8A-8B describes example operations that may be utilized by an audio remixing tool to construct dynamic audio presentation 800 for playing a composite audio program on real-world loudspeakers in a real-world venue. Those skilled in the art will recognize that these operations may be performed independently or in any combination to construct other dynamic audio presentations for playing other composite audio programs on real-world loudspeakers in real-world venues without departing from the spirit and scope of this disclosure.

[0073] 8A , an audio remixer can access an electronic library of audio voices 702, having audio voices 704.1-704.m present in a composite audio program, in a manner substantially similar to that described herein. The audio remixing tool can then assign the audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k in a virtual venue 600 and build a dynamic audio presentation 800 for playing the composite audio program on real-world loudspeakers in a real-world venue. As described herein, the virtual venue 600 can include a proscenium virtual loudspeaker system 604, effects-enhanced virtual array real-world loudspeaker systems 608.1-608.l, and / or environmental virtual array real-world loudspeaker systems 610.1-610.m, having virtual loudspeakers 606.1-606.z. As illustrated in FIG. 8A, the audio remixing tool can assign audio voices 704.1 generated by a microphone to virtual loudspeaker 606.4 from within the proscenium virtual loudspeaker system 604, audio voices 704.2 generated by an acoustic guitar, an electric guitar, and a bass guitar to virtual loudspeaker 606.7 from within the proscenium virtual loudspeaker system 604, audio voices 704.3 generated by a bass drum to virtual loudspeaker 606.2 from within the proscenium virtual loudspeaker system 604, and / or audio voices 704.m generated by a snare drum to virtual loudspeaker 606.1 from within the proscenium virtual loudspeaker system 604, in a manner substantially similar to that described herein through 7F, to construct a dynamic audio presentation 800. Those skilled in the art will recognize that the assignment of audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k in virtual venue 600 as illustrated in FIG. 8A is for illustrative purposes and is not intended to be limiting.Those skilled in the art will recognize that audio remixing tools may assign audio voices 704.1-704.m to other virtual loudspeakers 602.1-602.k in the virtual venue 600 to build a dynamic audio presentation 800 without departing from the spirit and scope of the present disclosure.

[0074] After assigning the audio voices 704.1-704.m to the virtual loudspeakers 602.1-602.k, the audio remixing tool can perform simple spatial movement of the audio voices 704.1-704.m within the three-dimensional space of the virtual venue 600 according to one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m. As described herein, the audio remixing tool can access one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m, such as, for example, spatial movement of the simple instruments and / or collections of instruments that produced the audio voices 704.1-704.m. Alternatively, or in addition, the audio remixing tool may analyze audio sounds 704.1-704.m, for example, audio sound 704.m produced by a snare drum as illustrated in FIG. 8A, and identify one or more characteristics, parameters, and / or attributes of audio sounds 704.1-704.m in a manner substantially similar to that described herein.

[0075] 8A , one or more characteristics, parameters, and / or attributes of audio sounds 704.1-704.m may indicate that the audio remixing tool will perform a simple spatial translation 810 within the three-dimensional space of virtual venue 600. In some embodiments, simple spatial translation 810 may include, for example, a spatial translation of audio sounds 704.m generated by a snare drum within the three-dimensional space of virtual venue 600 from virtual loudspeakers 606.1-606.z to effect-augmented virtual array real-world loudspeaker systems 608.1-608.1. In some embodiments, one or more characteristics, parameters, and / or attributes of audio sounds 704.1-704.m may indicate that the audio remixing tool will perform a simple spatial translation 810 within the three-dimensional space of virtual venue 600. In some embodiments, one or more characteristics, parameters, and / or attributes of audio sounds 704.1-704.m may indicate that the audio sounds 704.m generated by a snare drum will translate from virtual loudspeakers 606.1-606.z to effect-augmented virtual array real-world loudspeaker systems 608.1-608.1. move 8A to the effects-enhanced virtual array real-world loudspeaker systems 608.1-608.1. Those skilled in the art will recognize that the simple spatial translation 810 from the virtual loudspeakers 606.1-606.z to the effects-enhanced virtual array real-world loudspeaker systems 608.1-608.1 as illustrated in FIG. 8A is for illustrative purposes and is not intended to be limiting. Those skilled in the art will recognize that an audio remixing tool may perform other simple spatial translations for other audio voices from among the audio voices 704.1-704.m in a manner substantially similar to the simple spatial translation 810 without departing from the spirit and scope of the present disclosure.

[0076] In some embodiments, a simple spatial translation 810 is performed over a time t move6.1-608.l。 Alternatively, or in addition, the simple spatial movement 810 may represent a gradual spatial movement in the three-dimensional space of the virtual venue 600 from the virtual loudspeakers 606.1-606.z to the effect-augmented virtual array real-world loudspeaker systems 608.1-608.l, and then a snap spatial movement to the effect-augmented virtual array real-world loudspeaker systems 608.1-608.l. In some embodiments, the snap spatial movement may occur in response to an event, such as, for example, a period in the audio sound 704.m during which the snare drum does not produce audio sound, also referred to as a gap. For example, the snap spatial movement can be an instantaneous snap between virtual loudspeakers. As another example, the snap spatial movement can be a slow, gradual snap away from a virtual loudspeaker.

[0077] 8B graphically illustrates complex spatial movement operations 820 that may be performed by an audio remixing tool to assign audio sounds produced by simple instruments, such as percussion, wind, string, and / or electronic instruments, and / or collections of instruments, such as instruments from the same and / or different instrument classes, such as percussion, wind, string, and / or electronic instruments, to virtual loudspeakers 602.1-602.k in virtual venue 600. As described herein, the audio remixing tool may isolate attack transients from among audio sounds 704.1-704.m and decay transients from among audio sounds 704.1-704.m in a manner substantially similar to that described herein.

[0078] As shown in Figure 8B, the audio remixing tool can isolate, in a manner substantially similar to that described herein, the attack transient from within the audio audio 704.m produced by the snare drum, referred to as attack transient audio audio 752 in Figure 8A, and the decay transient from within the audio audio 704.m produced by the snare drum, referred to as decay transient audio audio 754 in Figure 8A. After isolating the attack and decay transients, the audio remixing tool can identify attack-decay real-world loudspeaker pairings from within the virtual loudspeakers 602.1-602.k in the virtual venue 600, in a manner substantially similar to that described herein. In the exemplary embodiment illustrated in FIG. 8B , the audio remixing tool can assign attack transients generated by simple instruments and / or collections of instruments, such as attack transient audio voice 752, to a virtual loudspeaker 756 from among the virtual loudspeakers 606.1-606.z, and the audio remixing tool can similarly assign decay transients generated by simple instruments and / or collections of instruments, such as decay transient audio voice 754, to a virtual loudspeaker 756 from among the virtual loudspeakers 606.1-606.z. Although the attack and decay audio sounds produced by simple instruments and / or collections of instruments are described as being assigned to virtual loudspeaker 756 in FIG. 8B, those skilled in the art will recognize that the attack and / or decay audio sounds may be assigned to any of virtual loudspeakers 602.1-602.k in virtual venue 600 in a manner substantially similar to that described in FIG. 8B without departing from the spirit and scope of the present disclosure.

[0079] After assigning the attack and decay audio voices, the audio remixing tool can perform complex spatial movement of the audio voices 704.1-704.m within the three-dimensional space of the virtual venue 600 according to one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m. As described herein, the audio remixing tool can access one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m, such as, for example, the spatial movement of the simple instruments and / or collections of instruments that produced the audio voices 704.1-704.m. Alternatively, or in addition, the audio remixing tool may analyze audio sounds 704.1-704.m, for example, audio sound 704.m produced by a snare drum as illustrated in FIG. 8A, and identify one or more characteristics, parameters, and / or attributes of audio sounds 704.1-704.m in a manner substantially similar to that described herein.

[0080] 8B , one or more characteristics, parameters, and / or attributes of audio sounds 704.1-704.m may indicate that the audio remixing tool performs a complex spatial movement 820 within the three-dimensional space of the virtual venue 600. In some embodiments, the complex spatial movement 820 may include, for example, a simple spatial movement 812 of an attack transient audio sound 752 generated by a snare drum within the three-dimensional space of the virtual venue 600 from virtual loudspeakers 606.1-606.z to effects-augmented virtual array real-world loudspeaker systems 608.1-608.1, and a simple spatial movement 814 of an attack transient audio sound 752 generated by a snare drum within the three-dimensional space of the virtual venue 600 from virtual loudspeakers 606.1-606.z to effects-augmented virtual array real-world loudspeaker systems 608.1-608.1. Those skilled in the art will recognize that complex spatial translation 820 from virtual loudspeakers 606.1-606.z to effects-augmented virtual array real-world loudspeaker system 608.1 as illustrated in FIG. 8B is for illustrative purposes and is not intended to be limiting. Those skilled in the art will recognize that an audio remixing tool may perform other simple spatial translations for other audio sounds from among audio sounds 704.1-704.m in a manner substantially similar to complex spatial translation 820 without departing from the spirit and scope of the present disclosure. In some embodiments, simple spatial translation 812 and / or simple spatial translation 814 can be performed in a manner substantially similar to that described herein. In these embodiments, simple spatial translation 812 can be performed before, simultaneously with, or after simple spatial translation 814. Example Configuration of an Example Real-World Venue for Implementing an Example Audio Presentation

[0081] 9 diagrammatically illustrates a high-level graphical mapping of an exemplary virtual venue to an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. In the exemplary embodiment illustrated in FIG. 9, an audio remixing tool, such as audio remixing tool 154, can assign audio voices, such as audio voices 156.1-156.n, to virtual loudspeakers 602.1-602.k in virtual venue 600 and construct an audio presentation, such as static audio presentation 700 and / or dynamic audio presentation 800, for playing a composite audio program in a real-world venue, such as real-world venue 102. As described herein, the audio remixing tool can configure the real-world venue as defined in the audio presentation to play the composite audio program in the real-world venue.

[0082] 9 , the audio remixing tool can assign audio sounds to virtual loudspeakers 602.1-602.k within the three-dimensional space of the virtual venue 600 to construct an audio presentation in a manner substantially similar to that described herein. In some embodiments, the real-world venue can include audio equipment such as amplifiers, crossovers, equalizers, and / or mixers to route and / or signal condition the audio sounds for playback within the real-world venue. In these embodiments, the audio remixing tool 154 can generate audio control signals, such as audio control signals 158.1-158.i, to configure the audio equipment in the real-world venue to play the audio presentation on the real-world loudspeakers in the real-world venue. In these embodiments, the audio control signals can cause the audio equipment in the real-world venue to route and / or signal condition the audio sounds for playback through the real-world loudspeakers in the real-world venue as defined in the audio presentation.

[0083] 9 , the audio remixing tool can access real-world venue configuration information 902, generate audio control signals, such as audio control signals 158.1-158.i, and configure audio equipment in the real-world venue to play an audio presentation on real-world loudspeakers in the real-world venue. In some embodiments, the real-world venue configuration information 902 represents an organized collection of data, often referred to as a database. A database may include one or more data tables having data values ​​such as alphanumeric strings, integers, decimals, floating-point numbers, dates, times, binary values, Boolean values, and / or enumerations, to name a few. A database can be a column-oriented database, a relational database, a keystore database, a graph database, and / or a document store, to name a few. In some embodiments, the real-world venue configuration information 902 can include real-world loudspeaker configuration information 904.1-904.k corresponding to virtual loudspeakers 602.1-602.k installed in the three-dimensional space of the virtual venue 600. In some embodiments, the audio remixing tool can access the real-world loudspeaker information 904.1-904.k and generate audio control signals for playing audio voices assigned to the virtual loudspeakers 602.1-602.k on the real-world loudspeakers in the real-world venue.In these embodiments, the audio remixing tool accesses real-world loudspeaker information 904.1 and generates audio control signals for playing audio voices assigned to virtual loudspeaker 602.1 on real-world loudspeakers in the real-world venue associated with virtual loudspeaker 602.1, and accesses real-world loudspeaker information 904.2 and generates audio control signals for playing audio voices assigned to virtual loudspeaker 602.2 on real-world loudspeakers in the real-world venue associated with virtual loudspeaker 602.2. The real-world loudspeaker information 904.k may be used to generate an audio control signal, access real-world loudspeaker information 904.3, and generate an audio control signal for playing an audio sound assigned to virtual loudspeaker 602.3 on a real-world loudspeaker in the real-world venue associated with virtual loudspeaker 602.3, and access real-world loudspeaker information 904.k, and generate an audio control signal for playing an audio sound assigned to virtual loudspeaker 602.k on a real-world loudspeaker in the real-world venue associated with virtual loudspeaker 602.k. Example Operation of an Example Audio Remixing Tool

[0084] 10 illustrates a flowchart of an exemplary audio remixing tool according to some exemplary embodiments of the present disclosure. The present disclosure is not limited to this operational description. Rather, other operational control flows will be apparent to those skilled in the art and are within the scope and spirit of the present disclosure. The following discussion describes an operational control flow 1000 for intelligently constructing an audio presentation and playing a composite audio program, such as composite audio program 150, in a real-world venue, such as real-world venue 102. The operational control flow 1000 can be implemented, for example, by audio remixing tool 154. In some embodiments, the operational control flow 1000 can be executed by one or more computing devices, such as audio playback server 104.

[0085] At operation 1002, the operation control flow 1000 constructs an audio presentation for playing the composite audio program in the real-world venue. The operation control flow 1000 can utilize a software remixing toolkit to construct the audio presentation in a manner substantially similar to that described herein.

[0086] At operation 1004, operation control flow 1000 can configure the real-world venue as defined in the audio presentation to play the composite audio program within the real-world venue. Operation control flow 1000 can identify audio control signals, such as audio control signals 158.1-158.i, that configure the real-world venue to play the audio presentation through real-world loudspeakers within the real-world venue to play the composite audio program within the real-world venue, in a manner substantially similar to that described herein. Example Computing Devices That Can Be Utilized to Implement Electronic Devices in an Example Real-World Venues

[0087] 11 diagrammatically illustrates a simplified block diagram of a computing device that may be utilized to implement electronic devices within an exemplary real-world venue, in accordance with some embodiments of the present disclosure. The following discussion of FIG. 11 describes a computing device 1100 that may be used to implement the audio playback server 104.

[0088] In the embodiment illustrated in FIG. 11, computing device 1100 includes one or more processors 1102. In some embodiments, one or more processors 1102 can include or be any of a microprocessor, a graphics processing unit, or a digital signal processor, and their electronic processing equivalents, such as an application-specific integrated circuit ("ASIC") or a field-programmable gate array ("FPGA"). As used herein, the term "processor" refers to a tangible data and information processing device that physically transforms data and information using sequence transformations (also referred to as "computations"). Data and information can be physically represented by electrical, magnetic, optical, or acoustic signals that can be stored, accessed, transferred, combined, compared, or otherwise manipulated by the processor. The term "processor" can refer to single processors as well as multi-core systems or multi-processor arrays that include a graphics processing unit, a digital signal processor, a digital processor, or a combination of these elements. A processor can be an electronic device, for example, comprising digital logic circuitry (e.g., binary logic) or analog (e.g., operational amplifiers). The processor may also operate to support the performance of related operations within a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some of the operations may be performed by a collection of processors available in a distributed or remote system, which are accessible via a communications network (e.g., the Internet) and via one or more software interfaces (e.g., application program interfaces (APIs)).In some embodiments, computing device 1100 may include an operating system such as Microsoft Windows®, Sun Microsystems' Solaris®, Apple Computer's MacOS, Linux®, or UNIX®. In some embodiments, computing device 1100 may also include a basic input / output system (BIOS) and processor firmware. The operating system, BIOS, and firmware are used by one or more processors 1102 to control subsystems and interfaces coupled to one or more processors 1102. In some embodiments, one or more processors 1102 may include Pentium® and Itanium processors manufactured by Intel, Opteron and Athlon processors manufactured by Advanced Micro Devices, and ARM processors manufactured by ARM Holdings.

[0089] 11, computing device 1100 can include machine-readable media 1104. In some embodiments, machine-readable media 1104 can further include main random access memory (“RAM”) 1106, read-only memory (“ROM”) 1108, and / or file storage subsystem 1110. RAM 730 can store instructions and data during program execution, while ROM 732 can store fixed instructions. File storage subsystem 1110 provides persistent storage for program and data files and may include a hard disk drive, a floppy disk drive, a CD-ROM drive, an optical drive, a flash memory, or a removable media cartridge, in addition to an associated removable media.

[0090] The computing device 1100 may further include user interface input devices 1112 and user interface output devices 1114. The user interface input devices 1112 may include pointing devices such as an alphanumeric keyboard, keypad, mouse, trackball, touchpad, stylus, or graphics tablet, scanner, touchscreen integrated into a display, audio input devices such as a voice recognition system or microphone, eye gaze recognition, electroencephalogram pattern recognition, and other types of input devices, to name a few. The user interface input devices 1112 may be connected to the computing device 1100 by wire or wirelessly. Generally, the user interface input devices 1112 are intended to include all possible types of devices and methods for inputting information into the computing device 1100. The user interface input devices 1112 typically allow a user to identify objects, icons, text, and the like that appear on some type of user interface output device, e.g., a display subsystem. The user interface output devices 1120 may include non-visual displays such as a display subsystem, printer, fax machine, or audio output device. The display subsystem may include a flat panel device such as a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other device for producing a visible image, such as a virtual reality system. The display subsystem may also provide a non-visual display, such as via audio output or haptic output (e.g., vibration) devices. Generally, user interface output devices 1120 are intended to include all possible types of devices and methods for outputting information from computing device 1100.

[0091] The computing device 1100 may further include a network interface 1116 for providing an interface to outside networks, including an interface to a communications network 1118, via which it is coupled to corresponding interface devices in other computing devices or machines. The communications network 1118 may comprise many interconnected computing devices, machines, and communications links. These communications links may be wired, optical, wireless, or any other devices for communicating information. The communications network 1118 may be any suitable computer network, for example, a wide area network such as the Internet and / or a local area network such as Ethernet. The communications network 1118 may be wired and / or wireless, and the communications network may use encryption and decryption methods such as those available with virtual private networks. The communications network uses one or more communications interfaces that may receive data from and transmit data to other systems. Embodiments of the communication interface typically include an Ethernet card, a modem (e.g., telephone, satellite, cable, or ISDN), an (asynchronous) Digital Subscriber Line (DSL) unit, a Firewire interface, a USB interface, and the like. One or more communication protocols, such as HTTP, TCP / IP, RTP / RTSP, IPX, and / or UDP, can be used.

[0092] 11, one or more processors 1102, machine-readable media 1104, user interface input devices 1112, user interface output devices 1114, and / or network interface 1116 can be communicatively coupled to each other using a bus subsystem 1120. Although the bus subsystem 1120 is shown diagrammatically as a single bus, alternative embodiments of the bus subsystem may use a bus. For example, a RAM-based main memory can communicate directly with a file storage system using a direct memory access (“DMA”) system. (Conclusion)

[0093] The detailed description has referred to the accompanying figures to illustrate exemplary embodiments consistent with this disclosure. References to an "exemplary embodiment" in this disclosure indicate that the described exemplary embodiment may include a particular feature, structure, or characteristic, but not all exemplary embodiments necessarily include the particular feature, structure, or characteristic. Also, such phrases do not necessarily refer to the same exemplary embodiment. Furthermore, any feature, structure, or characteristic described in connection with an exemplary embodiment can be included independently or in any combination with features, structures, or characteristics of other exemplary embodiments, whether or not explicitly described.

[0094] The Detailed Description is not meant to be limiting. Rather, the scope of the present disclosure is defined solely by the following claims and their equivalents. It is understood that the Detailed Description section, and not the Abstract section, is intended to be used to interpret the claims. The Abstract section may describe one or more example embodiments, but not all, of the present disclosure, and is therefore not intended to limit the present disclosure and the following claims and their equivalents in any way.

[0095] The exemplary embodiments described within this disclosure are provided for illustrative purposes and are not intended to be limiting. Other exemplary embodiments are possible, and modifications can be made to the exemplary embodiments while remaining within the spirit and scope of this disclosure. This disclosure is described with the help of functional building blocks that illustrate implementations of defined functions and their relationships. The boundaries of these functional building blocks are arbitrarily defined herein for convenience of description. Alternative boundaries can be defined so long as the defined functions and their relationships are appropriately performed.

[0096] Embodiments of the present disclosure may be implemented in hardware, firmware, a software application, or any combination thereof. Embodiments of the present disclosure may also be implemented as instructions stored on a machine-readable medium, which may be read and executed by one or more processors. A machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing network). For example, a machine-readable medium may include non-transitory machine-readable media, such as read-only memory (ROM), random-access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, and the like. As another example, a machine-readable medium may include transitory machine-readable media, such as electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). Furthermore, firmware, software applications, routines, instructions may be described herein as performing certain actions. However, it should be understood that such description is for convenience only, and that such actions actually result from a computing device, processor, controller, or other device executing the firmware, software application, routines, instructions, etc.

[0097] The detailed description of the exemplary embodiments has fully revealed the general nature of the present disclosure, such that others, by applying the knowledge of those skilled in the art, may readily modify and / or adapt such exemplary embodiments for use without undue experimentation, without departing from the spirit and scope of the present disclosure. Such adaptations and modifications are therefore intended to be within the meaning and equivalents of the exemplary embodiments, based on the teaching and guidance presented herein. It is to be understood that the phraseology or terminology herein is for the purpose of description, not limitation, and thus will be interpreted by one of ordinary skill in the art in light of the teachings herein.

Claims

1. 1. An audio playback server for demixing a composite audio program into a plurality of audio voices, said audio playback server comprising: a memory configured to store an audio demixing tool; a processor configured to execute the audio demixing tool, the audio demixing tool, when executed by the processor, selecting a sample audio sound from among a plurality of sample sounds; identifying an audio voice from the plurality of audio voices that corresponds to the sample audio voice; decomposing said audio sounds from said composite audio program; a processor, the processor configured to: Equipped with the audio demixing tool further configures the processor to, when executed by the processor, iteratively select the sample audio voices, identify the audio voices, and iteratively decompose the audio voices; The audio playback server further configuring the processor such that, for each iteration, the audio demixing tool, when executed by the processor, selects a different sample audio sound from among the plurality of sample sounds.

2. 2. The audio playback server of claim 1, wherein the audio demixing tool, when executed by the processor, configures the processor to search the composite audio program for the audio sounds that match the sample audio sounds.

3. 3. The audio playback server of claim 2, wherein the audio demixing tool, when executed by the processor, configures the processor to perform a pattern recognition algorithm to match the sample audio voice to the audio voice.

4. 2. The audio playback server of claim 1, wherein the audio demixing tool, when executed by the processor, configures the processor to isolate the audio voices from the composite audio program and to decompose the audio voices from the composite audio program.

5. 5. The audio playback server of claim 4, wherein the audio demixing tool, when executed by the processor, configures the processor to subtract the audio voice from the composite audio program and isolate the audio voice from the composite audio program.

6. 10. The audio playback server of claim 1, wherein the audio demixing tool, when executed by the processor, further configures the processor to analyze the audio speech and identify one or more characteristics, parameters, or attributes corresponding to the audio speech.

7. 7. The audio playback server of claim 6, wherein the one or more characteristics, parameters, or attributes corresponding to the audio voice comprise the pitch, volume, timbre, frequency, amplitude, wavelength, or rate of the audio voice.

8. 1. A method for demixing a composite audio program into a plurality of audio voices, said method comprising: a computing device performing an iterative process for demixing the composite audio program into the plurality of audio voices. Including, The iterative process comprises: selecting a sample audio sound from among a plurality of sample sounds by the computing device; the computing device identifying an audio voice from among the plurality of audio voices that corresponds to the sample audio voice; said computing device decomposing said audio sounds from said composite audio program; Including, The method, wherein, for each iteration of the iterative process, the selecting selects a different sample audio sound from among the plurality of sample sounds.

9. 9. The method of claim 8, wherein said identifying comprises searching said composite audio program for said audio sounds that match said sample audio sounds.

10. The method of claim 9 , wherein the searching includes running a pattern recognition algorithm to match the sample audio speech to the audio speech.

11. 9. The method of claim 8, wherein said decomposing comprises isolating said audio voice from said composite audio program and decomposing said audio voice from said composite audio program.

12. The method of claim 11 , wherein said isolating comprises subtracting said audio voice from said composite audio program and isolating said audio voice from said composite audio program.

13. The method of claim 8 , further comprising analyzing the audio speech to identify one or more characteristics, parameters, or attributes corresponding to the audio speech.

14. The method of claim 13 , wherein the one or more characteristics, parameters, or attributes corresponding to the audio sound comprise pitch, volume, timbre, frequency, amplitude, wavelength, or rate of the audio sound.

15. A venue for playing a composite audio program, said venue comprising: An audio playback server, the audio playback server comprising: selecting a sample audio sound from among a plurality of sample sounds; identifying an audio speech from among a plurality of audio speeches that corresponds to the sample audio speech; decomposing said audio sounds from said composite audio program; configured to: the audio playback server is further configured to iteratively select the sample audio sounds, identify the audio sounds, and iteratively decompose the audio sounds; an audio playback server, wherein for each iteration, the audio playback server is further configured to select a different sample audio sound from among the plurality of sample sounds; a plurality of loudspeakers configured to reproduce the plurality of audio sounds; A venue equipped with:

16. 16. The venue of claim 15, wherein the audio playback server is configured to search the composite audio program for the audio sounds that match the sample audio sounds.

17. 17. The venue of claim 16, wherein the audio playback server is configured to run a pattern recognition algorithm to match the sample audio voice to the audio voice.

18. 16. The venue of claim 15, wherein the audio playback server is configured to isolate and decouple the audio sounds from the composite audio program.

19. 20. The venue of claim 18, wherein the audio playback server is configured to subtract and isolate the audio sounds from the composite audio program.

20. 16. The venue of claim 15, wherein the audio playback server is configured to analyze the audio sounds and identify one or more characteristics, parameters, or attributes corresponding to the audio sounds.

21. 1. An audio playback server for remixing a plurality of audio voices of a composite audio program for playback in a real-world venue, said audio playback server comprising: a memory configured to store an audio remixing tool; a processor configured to execute the audio remixing tool, the audio remixing tool, when executed by the processor, accessing a virtual venue corresponding to the real-world venue, the virtual venue including a plurality of virtual loudspeakers corresponding to a plurality of real-world loudspeakers in the venue; accessing an ensemble bounding volume that defines a maximum spatial distance between a first audio voice and a second audio voice of an audio voice ensemble from among the plurality of audio voices; assigning the first audio voice to a first virtual loudspeaker from among the plurality of virtual loudspeakers; assigning the second audio sound to a second virtual loudspeaker from among the plurality of virtual loudspeakers that is less than the maximum spatial distance from the first audio sound; identifying a plurality of audio control signals within the real-world venue and configuring the real-world venue to play the first audio sound on a first real-world loudspeaker from among the plurality of real-world loudspeakers corresponding to the first virtual loudspeaker and to play the second audio sound on a second real-world loudspeaker from the plurality of real-world loudspeakers corresponding to the second virtual loudspeaker; a processor, the processor configured to: An audio playback server comprising:

22. 22. The audio playback server of claim 21, wherein the audio remixing tool, when executed by the processor, further configures the processor to access an electronic library of audio voices comprising the plurality of audio voices parsed from the composite audio program.

23. The audio remixing tool, when executed by the processor, logically grouping the first audio voice and the second audio voice to form the audio voice ensemble; positioning the first audio sound and the second audio sound at the spatial distance within the ensemble bounding volume; determining an audible time of flight for the first audio voice and the second audio voice to arrive at a location within the virtual venue; further configuring the processor to: the audio remixing tool further configures the processor to, when executed by the processor, iteratively position the first audio voice and the second audio voice, iteratively determine the audible time of flight, and form the ensemble bounding volume; 22. The audio playback server of claim 21 , wherein the audio remixing tool, when executed by the processor, for each iteration further configures the processor to adjust the spatial distance to determine the maximum spatial distance until a difference between the audible times of flight exceeds an audible time-of-flight threshold.

24. 22. The audio playback server of claim 21, wherein the audio remixing tool, when executed by the processor, further configures the processor to move the second audio voice from the second virtual loudspeaker to a third virtual loudspeaker that is less than the maximum spatial distance from the first audio voice during the composite audio program.

25. 25. The audio playback server of claim 24, wherein the movement of the second audio voice from the second virtual loudspeaker to the third virtual loudspeaker comprises a snap spatial movement of the second audio voice.

26. The audio remixing tool, when executed by the processor, isolating an attack transient of the first audio sound and a decay transient of the first audio sound from the first audio sound; assigning the attack transient of the first audio sound to the first virtual loudspeaker and the decay transient of the first audio sound to a third virtual loudspeaker from among the plurality of virtual loudspeakers that is less than the maximum spatial distance from the first audio sound; 22. The audio playback server of claim 21, further configured to:

27. The audio remixing tool, when executed by the processor, moving the attack transient of the first audio sound from the first virtual loudspeaker to a fourth virtual loudspeaker that is less than the maximum spatial distance from the first audio sound during the composite audio program; or moving the decay transients of the first audio voice from the third virtual loudspeaker to the fourth virtual loudspeaker during the composite audio program.

27. The audio playback server of claim 26, further configured to:

28. 1. A method for remixing multiple audio voices of a composite audio program for playback in a real-world venue, the method comprising: accessing, by a computing device, a virtual venue corresponding to the real-world venue, the virtual venue including a plurality of virtual loudspeakers corresponding to a plurality of real-world loudspeakers in the venue; accessing, by the computing device, an ensemble bounding volume that defines a maximum spatial distance between a first audio voice and a second audio voice of an audio voice ensemble from among the plurality of audio voices; the computing device assigning the first audio voice to a first virtual loudspeaker from among the plurality of virtual loudspeakers; the computing device assigning the second audio sound to a second virtual loudspeaker from among the plurality of virtual loudspeakers that is less than the maximum spatial distance from the first audio sound; the computing device identifying a plurality of audio control signals within the real-world venue and configuring the real-world venue to play the first audio sound on a first real-world loudspeaker from among the plurality of real-world loudspeakers corresponding to the first virtual loudspeaker and to play the second audio sound on a second real-world loudspeaker from among the plurality of real-world loudspeakers corresponding to the second virtual loudspeaker; A method comprising:

29. 30. The method of claim 28, further comprising the computing device accessing an electronic library of audio sounds comprising the plurality of audio sounds parsed from the composite audio program.

30. the computing device logically grouping the first audio voice and the second audio voice to form the audio voice ensemble; the computing device positioning the first audio sound and the second audio sound at the spatial distance within the ensemble bounding volume; the computing device determining an audible time of flight for the first audio sound and the second audio sound to arrive at a location within the virtual venue; repeating the locating iteratively to determine the audible time of flight and form the ensemble bounding volume; further comprising 30. The method of claim 28, wherein for each iteration, the spatial distance is adjusted to determine the maximum spatial distance until the difference between the audible times of flight exceeds an audible time-of-flight threshold.

31. 29. The method of claim 28, further comprising the computing device moving the second audio sound from the second virtual loudspeaker to a third virtual loudspeaker that is less than the maximum spatial distance from the first audio sound during the composite audio program.

32. 32. The method of claim 31 , wherein the movement of the second audio sound from the second virtual loudspeaker to the third virtual loudspeaker comprises a snap spatial movement of the second audio sound.

33. the computing device isolating from the first audio sound an attack transient of the first audio sound and a decay transient of the first audio sound; the computing device assigning the attack transient of the first audio sound to the first virtual loudspeaker and the decay transient of the first audio sound to a third virtual loudspeaker from among the plurality of virtual loudspeakers that is less than the maximum spatial distance from the first audio sound; 30. The method of claim 28, further comprising:

34. the computing device moving the attack transient of the first audio sound from the first virtual loudspeaker to a fourth virtual loudspeaker that is less than the maximum spatial distance from the first audio sound during the composite audio program; or the computing device moving the decay transients of the first audio voice from the third virtual loudspeaker to the fourth virtual loudspeaker during the composite audio program.

34. The method of claim 33, further comprising:

35. A real-world venue for playing a composite audio program, said real-world venue comprising: a plurality of real-world loudspeakers in the real-world venue configured to play a plurality of audio sounds of the composite audio program; An audio playback server, the audio playback server comprising: accessing a virtual venue corresponding to the real-world venue, the virtual venue including a plurality of virtual loudspeakers corresponding to the plurality of real-world loudspeakers; accessing an ensemble bounding volume that defines a maximum spatial distance between a first audio voice and a second audio voice of an audio voice ensemble from among the plurality of audio voices; assigning the first audio voice to a first virtual loudspeaker from among the plurality of virtual loudspeakers; assigning the second audio sound to a second virtual loudspeaker from among the plurality of virtual loudspeakers that is less than the maximum spatial distance from the first audio sound; identifying a plurality of audio control signals within the real-world venue and configuring the real-world venue to play the first audio sound on a first real-world loudspeaker from among the plurality of real-world loudspeakers corresponding to the first virtual loudspeaker and to play the second audio sound on a second real-world loudspeaker from among the plurality of real-world loudspeakers corresponding to the second virtual loudspeaker; an audio playback server configured to: A real-world venue equipped with:

36. 36. The real-world venue of claim 35, wherein the audio playback server is further configured to access an electronic library of audio sounds comprising the plurality of audio sounds parsed from the composite audio program.

37. The audio playback server logically grouping the first audio voice and the second audio voice to form the audio voice ensemble; positioning the first audio sound and the second audio sound at the spatial distance within the ensemble bounding volume; determining an audible time of flight for the first audio voice and the second audio voice to arrive at a location within the virtual venue; further configured to: the audio playback server is further configured to iteratively position the first audio sound and the second audio sound, iteratively determine the audible time-of-flight, and form the ensemble bounding volume; 36. The real-world venue of claim 35, wherein, for each iteration, the audio playback server is further configured to adjust the spatial distance to determine the maximum spatial distance until a difference between the audible times of flight exceeds an audible time-of-flight threshold.

38. 36. The real-world venue of claim 35, wherein the audio playback server is further configured to move the second audio sound from the second virtual loudspeaker to a third virtual loudspeaker that is less than the maximum spatial distance from the first audio sound during the composite audio program.

39. 39. The real-world venue of claim 38, wherein the movement of the second audio sound from the second virtual loudspeaker to the third virtual loudspeaker comprises a snap spatial movement of the second audio sound.

40. The audio playback server isolating an attack transient of the first audio sound and a decay transient of the first audio sound from the first audio sound; assigning the attack transient of the first audio sound to the first virtual loudspeaker and the decay transient of the first audio sound to a third virtual loudspeaker from among the plurality of virtual loudspeakers that is less than the maximum spatial distance from the first audio sound; moving the attack transient of the first audio sound from the first virtual loudspeaker to a fourth virtual loudspeaker that is less than the maximum spatial distance from the first audio sound during the composite audio program; or moving the decay transients of the first audio voice from the third virtual loudspeaker to the fourth virtual loudspeaker during the composite audio program.

36. The real world venue of claim 35, further configured to: