Demixing and remixing of composite audio programs for playback within the venue.
The audio playback server addresses comb filtering and latency issues by demixing and remixing audio programs using pattern recognition and spatial audio management, enhancing audio synchronization and immersion in real-world venues.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SPHERE ENTERTAINMENT GROUP LLC
- Filing Date
- 2025-09-04
- Publication Date
- 2026-06-01
AI Technical Summary
The comb filtering effect and audio latency issues in real-world venues cause undesirable audio phenomena such as metallic sounds, robot-like effects, and synchronization problems between multiple audio tracks, which are particularly noticeable when time delays exceed 10-30 milliseconds, affecting live performances and audio productions.
An audio playback server demixes and remixes composite audio programs by using pattern recognition algorithms to isolate and analyze audio characteristics, then intelligently assigns audio signals to loudspeakers within a venue to mitigate comb filtering and latency, employing iterative processes and ensemble boundary volume adjustments.
The solution effectively decomposes and reconstructs audio signals to minimize comb filtering and latency, ensuring synchronized and immersive audio experiences across different venues.
Smart Images

Figure 2026089658000001_ABST
Abstract
Description
[Background technology]
[0001] (Cross-reference of related applications) This application claims the benefit of U.S. Provisional Patent Application No. 63 / 449,797, filed on 3 March 2023, which is incorporated herein by reference in its entirety.
[0002] The comb filtering effect describes a phenomenon in a real-world venue that occurs when audio sound and its reflections arrive at a location within the venue at different times. Audio sound can be reflected when it comes into contact with a hard surface in the real-world venue, such as the floor, walls, windows, or equivalent. Because reflections travel longer distances than the audio sound, they may arrive at their location later than the audio sound. Often, certain frequencies of the audio sound are amplified or attenuated by the superposition of its reflection onto itself, causing the comb filtering effect. This superposition can cause some cancellation and amplification of the audio spectrum, which can produce subjectively metallic-sounding audio sound, also known as metallic sound. Generally, when the time delay between the audio sound and its reflection is about 20 to 30 milliseconds, the human ear may, undesirably, perceive the audio sound and its reflection as separate signals. For example, when the time delay between an audio sound and its reflection is about 50 milliseconds, the human ear begins to perceive the reflection as an echo of the audio sound, and this becomes even more apparent at about 100 milliseconds. However, the comb filtering effect can also be present when the time delay between an audio sound and its reflection is about 12 to 15 milliseconds. For example, the comb filtering effect can alter the timbre of a sound at a delay of about 12 milliseconds between the audio sound and its reflection. The comb filtering effect can also make an audio sound more "robot-like" by increasing the time delay between the audio sound and its reflection. In some situations, these comb filtering effects are not noticeable when the time delay between the audio sound and its reflection is about 3 to 10 milliseconds.
[0003] Audio latency refers to the time delay between when an audio sound is produced and when it arrives at a specific location within a real-world venue. In the context of a real-world venue, audio latency can be influenced by various factors, including the distance between the audio source and its location within the venue, the acoustic effects of the venue, the processing and transmission of the audio sound, and the performance of any digital processing or effects applied to the audio sound. Often, multiple audio sounds produced by multiple real-world loudspeakers within a venue may arrive at a single location within the venue at different times. This can cause audio latency between these audio sounds within the venue. Often, when the audio latency between multiple audio sounds is between approximately 10 (10) milliseconds and less than approximately 12 (12) milliseconds, the time delay between the multiple audio sounds is more likely to go unnoticed. However, when the audio latency between multiple audio sounds is between approximately 20 (20) milliseconds and approximately 30 (30) milliseconds, the human ear may perceive these audio sounds as separate signals. For example, audio latency can cause fading, reverberation, or even a lack of synchronization between multiple audio tracks. In live performance or audio production, managing audio latency between multiple audio tracks is crucial for maintaining consistent and synchronized sound. [Overview of the Initiative] [Means for solving the problem]
[0004] Detailed explanation The following disclosure provides many different embodiments or examples for implementing different features of the subject matter provided. Specific examples of components and arrangements are described herein for the sake of brevity of this disclosure. These are, of course, examples only and are not intended to be limiting. Aspects of this disclosure are best understood from the following detailed description when read carefully together with the accompanying figures. This disclosure may repeat reference numbers and / or letters in various embodiments. This repetition does not in itself determine the relationships between the various embodiments and / or configurations discussed. Note that, in accordance with standard practice in this industry, features are not drawn to scale. In fact, the dimensions of features may be increased or decreased as appropriate for the sake of clarity of discussion.
[0005] The following disclosure may include spatially relative terms such as “directly below,” “downward,” “below,” “up,” “top,” “upper,” and equivalents to facilitate explanation of the relationships between elements or features as illustrated in the figures herein. These spatially relative terms are intended to encompass different orientations relating to different embodiments or examples depicted in the figures. Different embodiments or examples may be oriented differently (rotated by 90 degrees or other orientations), and the spatially relative terms included herein may be interpreted accordingly. The following disclosure may also include the terms “about,” “approximately,” or “substantially” to indicate that, based on the particular art, the value of a given quantity may vary. Based on the art, the terms “about” or “substantially” may indicate a value of a given quantity that varies, for example, within 1 to 15% of the value (e.g., ±1%, ±2%, ±5%, ±10%, or ±15% of the value). (Overview)
[0006] The systems, methods, and apparatus can reproduce a composite audio program associated with an event held by a real-world venue. These systems, methods, and apparatus can seamlessly decompose the composite audio program into multiple audio sounds that can be collectively reproduced by the real-world venue. As part of this decomposition, these systems, methods, and apparatus can analyze the audio sounds and identify one or more characteristics, parameters, and / or attributes of these audio sounds. These systems, methods, and apparatus can intelligently construct an audio presentation from these multiple audio sounds and reproduce the composite audio program within the real-world venue. As part of this construction, these systems, methods, and apparatus construct the audio presentation based on one or more characteristics, parameters, and / or attributes of these audio sounds. After constructing the audio presentation, these systems, methods, and apparatus can configure the real-world venue as defined in the audio presentation to reproduce the composite audio program within the real-world venue. As part of this construction, these systems, methods, and apparatus can identify audio control signals that configure the real-world venue to reproduce the audio presentation through real-world loudspeakers within the real-world venue. The present invention provides, for example, the following: (Item 1) An audio playback server for demixing a composite audio program into multiple audio sources, wherein the audio playback server is Memory configured to store audio demixing tools, A processor configured to execute the aforementioned audio demixing tool, wherein the audio demixing tool, when executed by the processor, Select a sample audio from among several sample audios, From among the multiple audio sounds, identify the audio sound corresponding to the sample audio sound, Decomposing the audio sound from the aforementioned composite audio program The processor is configured to perform the following: Equipped with, The audio demixing tool, when executed by the processor, further configures the processor to iteratively select the sample audio sounds, identify the audio sounds, and iteratively demix the audio sounds. For each iteration, the audio demixing tool further configures the processor to select another sample audio from the plurality of sample audios, as executed by the processor, in an audio playback server. (Item 2) The audio playback server according to item 1, wherein the audio demixing tool, when executed by the processor, configures the processor to search for the composite audio program with respect to the audio sound that matches the sample audio sound. (Item 3) The audio playback server described in item 2, wherein the audio demixing tool, when executed by the processor, configures the processor to execute a pattern recognition algorithm and match the sample audio to the audio. (Item 4) The audio playback server according to item 1, wherein the audio demixing tool, when executed by the processor, configures the processor to isolate the audio sound from the composite audio program and decompose the audio sound from the composite audio program. (Item 5) The audio playback server according to item 4, wherein the audio demixing tool, when executed by the processor, configures the processor to subtract the audio from the composite audio program and isolate the audio from the composite audio program. (Item 6) The audio playback server according to item 1, wherein the audio demixing tool, when executed by the processor, further configures the processor to analyze the audio and identify one or more characteristics, parameters, or attributes corresponding to the audio. (Item 7) The audio playback server according to item 6, wherein the one or more characteristics, parameters, or attributes corresponding to the audio sound include the pitch, volume, timbre, frequency, amplitude, wavelength, or velocity of the audio sound. (Item 8) A method for demixing a composite audio program into multiple audio sounds, wherein the method is The computing device performs an iterative process to demix the composite audio program into the plurality of audio sounds. Includes, The aforementioned iterative process is The computing device selects a sample audio from among multiple sample audios, The computing device identifies the audio corresponding to the sample audio from among the plurality of audio sounds, The computing device decomposes the audio sound from the composite audio program. Includes, For each iteration of the aforementioned iterative process, the selection is a method of selecting another sample audio from among the multiple sample audios. (Item 9) The method of item 8, wherein the identification includes searching for the composite audio program with respect to the audio sound that matches the sample audio sound. (Item 10) The method according to item 9, wherein the search includes executing a pattern recognition algorithm and matching the sample audio to the audio. (Item 11) The method of item 8, wherein the decomposition includes isolating the audio sound from the composite audio program and decomposing the audio sound from the composite audio program. (Item 12) The method according to item 11, wherein the isolation includes subtracting the audio sound from the composite audio program and isolating the audio sound from the composite audio program. (Item 13) The method according to item 8, further comprising analyzing the audio sound and identifying one or more characteristics, parameters, or attributes corresponding to the audio sound. (Item 14) The method according to item 13, wherein the one or more characteristics, parameters, or attributes corresponding to the audio sound include the pitch, volume, timbre, frequency, amplitude, wavelength, or velocity of the audio sound. (Item 15) A venue for playing a composite audio program, the venue is An audio playback server, wherein the audio playback server is Select a sample audio from among several sample audios, Identifying the audio corresponding to the sample audio from among multiple audio tracks, Decomposing the audio sound from the aforementioned composite audio program It is configured to do the following: The audio playback server is further configured to iteratively select the sample audio, identify the audio, and iteratively decompose the audio. In each iteration, the audio playback server is further configured to select another sample audio from the plurality of sample audios, Multiple loudspeakers configured to play multiple audio sounds and A venue equipped with these features. (Item 16) The venue according to item 15, wherein the audio playback server is configured to search for the composite audio program with respect to the audio that matches the sample audio. (Item 17) The venue described in item 16, wherein the audio playback server is configured to execute a pattern recognition algorithm and match the sample audio to the audio. (Item 18) The venue according to item 15, wherein the audio playback server is configured to isolate the audio sound from the composite audio program and to decompose the audio sound from the composite audio program. (Item 19) The venue according to item 18, wherein the audio playback server is configured to subtract the audio sound from the composite audio program and isolate the audio sound from the composite audio program. (Item 20) The venue described in item 15, wherein the audio playback server is configured to analyze the audio and identify one or more characteristics, parameters, or attributes corresponding to the audio. (Item 21) An audio playback server for remixing multiple audio sounds of a composite audio program for playback within a real-world venue, wherein the audio playback server is Memory configured to store audio remixing tools, A processor configured to execute the aforementioned audio remixing tool, wherein the audio remixing tool, when executed by the processor, Accessing a virtual venue corresponding to the aforementioned real-world venue, wherein the virtual venue includes a plurality of virtual loudspeakers corresponding to a plurality of real-world loudspeakers within the venue. Accessing the ensemble boundary volume that defines the maximum spatial distance between the first audio and the second audio of an audio ensemble selected from the aforementioned plurality of audio sounds, Assigning the first audio signal to the first virtual loudspeaker from among the plurality of virtual loudspeakers, Assigning the second audio to a second virtual loudspeaker from among the plurality of virtual loudspeakers that are less than the maximum spatial distance from the first audio, The real-world venue is configured such that multiple audio control signals within the real-world venue are identified, the first audio sound is played in the first real-world loudspeaker from among the multiple real-world loudspeakers corresponding to the first virtual loudspeaker, and the second audio sound is played in the second real-world loudspeaker from among the multiple real-world loudspeakers corresponding to the second virtual loudspeaker. The processor is configured to perform the following: An audio playback server equipped with the following features. (Item 22) The audio playback server according to item 21, wherein the audio remixing tool, when executed by the processor, further configures the processor to access an electronic library of audio sounds having the plurality of audio sounds decomposed from the composite audio program. (Item 23) When the aforementioned audio remixing tool is executed by the aforementioned processor, The first audio and the second audio are logically grouped to form the audio ensemble, Positioning the first audio sound and the second audio sound within the spatial distance of the ensemble boundary volume, Determining the audible flight time until the first audio signal and the second audio signal arrive at a certain location within the virtual venue. The processor is further configured to perform the following: The audio remixing tool, when executed by the processor, further configures the processor to iteratively position the first audio and the second audio, iteratively determine the audible flight time, and form the ensemble boundary volume. The audio playback server according to item 21, wherein, with each iteration, the audio remixing tool, when executed by the processor, further configures the processor to adjust the spatial distances in order to determine the maximum spatial distance, such that the difference between the audible times of flight exceeds an audible time of flight threshold. (Item 24) The audio playback server according to item 21, wherein the audio remixing tool, when executed by the processor, further configures the processor to move the second audio sound from the second virtual loudspeaker to a third virtual loudspeaker that is less than the maximum spatial distance from the first audio sound during the composite audio program. (Item 25) The movement of the second audio sound from the second virtual loudspeaker to the third virtual loudspeaker is an audio playback server according to item 24, comprising a snap spatial movement of the second audio sound. (Item 26) When the aforementioned audio remixing tool is executed by the aforementioned processor, To isolate the attack transient and decay transient of the first audio from the first audio, Assigning the attack transient of the first audio sound to the first virtual loudspeaker, and assigning the decay transient of the first audio sound to a third virtual loudspeaker from among the plurality of virtual loudspeakers that are less than the maximum spatial distance from the first audio sound. The audio playback server according to item 21, further configured to perform the above processor. (Item 27) When the aforementioned audio remixing tool is executed by the aforementioned processor, During the composite audio program, the attack transient of the first audio sound is moved from the first virtual loudspeaker to a fourth virtual loudspeaker that is less than the maximum spatial distance from the first audio sound, or During the composite audio program, the decay transient of the first audio sound is moved from the third virtual loudspeaker to the fourth virtual loudspeaker. The audio playback server according to item 26, further configured to perform the above processor. (Item 28) A method for remixing multiple audio sounds of a composite audio program for playback in a real-world venue, wherein the method is: The computing device accesses a virtual venue corresponding to the real-world venue, wherein the virtual venue includes a plurality of virtual loudspeakers corresponding to a plurality of real-world loudspeakers within the venue. The computing device accesses an ensemble boundary volume that defines the maximum spatial distance between the first audio and the second audio of an audio ensemble selected from the plurality of audio sounds, The computing device assigns the first audio signal to the first virtual loudspeaker from among the plurality of virtual loudspeakers, The computing device assigns the second audio to a second virtual loudspeaker from among the plurality of virtual loudspeakers that are less than the maximum spatial distance from the first audio, The computing device identifies a plurality of audio control signals within the real-world venue, and configures the real-world venue to play the first audio sound in the first real-world loudspeaker from among the plurality of real-world loudspeakers corresponding to the first virtual loudspeaker, and to play the second audio sound in the second real-world loudspeaker from among the plurality of real-world loudspeakers corresponding to the second virtual loudspeaker. Methods that include... (Item 29) The method according to item 28, further comprising the computing device accessing an electronic library of audio sounds having the plurality of audio sounds decomposed from the composite audio program. (Item 30) The computing device logically groups the first audio and the second audio to form the audio ensemble. The computing device positions the first audio and the second audio within the spatial distance of the ensemble boundary volume, The computing device determines the audible flight time until the first audio and the second audio arrive at a certain location within the virtual venue. The above positioning is repeated iteratively to determine the audible flight time and to form the ensemble boundary volume. It further includes, The method of item 28, wherein, with each iteration, the spatial distance is adjusted until the difference between the audible flight times exceeds an audible flight time threshold in order to determine the maximum spatial distance. (Item 31) The method of item 28, further comprising the computing device moving the second audio sound from the second virtual loudspeaker to a third virtual loudspeaker which is less than the maximum spatial distance from the first audio sound during the composite audio program. (Item 32) The method according to item 31, wherein the movement of the second audio sound from the second virtual loudspeaker to the third virtual loudspeaker comprises a snap spatial movement of the second audio sound. (Item 33) The computing device isolates the attack transient and decay transient of the first audio from the first audio. The computing device assigns the attack transient of the first audio sound to the first virtual loudspeaker, and assigns the decay transient of the first audio sound to a third virtual loudspeaker from among the plurality of virtual loudspeakers that is less than the maximum spatial distance from the first audio sound. The method described in item 28, further including the method described in item 28. (Item 34) The computing device, during the composite audio program, moves the attack transient of the first audio sound from the first virtual loudspeaker to a fourth virtual loudspeaker that is less than the maximum spatial distance from the first audio sound, or The computing device moves the decay transient of the first audio sound from the third virtual loudspeaker to the fourth virtual loudspeaker during the composite audio program. The method described in item 33, further including the method described in item 33. (Item 35) A real-world venue for playing a composite audio program, wherein the real-world venue is Multiple real-world loudspeakers in the real-world venue, configured to play multiple audio sounds from the composite audio program, An audio playback server, wherein the audio playback server is Accessing a virtual venue corresponding to the aforementioned real-world venue, wherein the virtual venue includes a plurality of virtual loudspeakers corresponding to the plurality of real-world loudspeakers, Accessing the ensemble boundary volume that defines the maximum spatial distance between the first audio and the second audio of an audio ensemble selected from the aforementioned plurality of audio sounds, Assigning the first audio signal to the first virtual loudspeaker from among the plurality of virtual loudspeakers, Assigning the second audio to a second virtual loudspeaker from among the plurality of virtual loudspeakers that are less than the maximum spatial distance from the first audio, Identifying multiple audio control signals within the real-world venue, configuring the real-world venue to reproduce the first audio sound in the first real-world loudspeaker from among the multiple real-world loudspeakers corresponding to the first virtual loudspeaker, and to reproduce the second audio sound in the second real-world loudspeaker from among the multiple real-world loudspeakers corresponding to the second virtual loudspeaker. An audio playback server and A real-world venue equipped with these features. (Item 36) The audio playback server is further configured to access an electronic library of audio sounds having the plurality of audio sounds decomposed from the composite audio program, as described in item 35, for the real-world venue. (Item 37) The aforementioned audio playback server The first audio and the second audio are logically grouped to form the audio ensemble, Positioning the first audio sound and the second audio sound within the spatial distance of the ensemble boundary volume, Determining the audible flight time until the first audio signal and the second audio signal arrive at a certain location within the virtual venue. It is further configured to do the following: The audio playback server is further configured to iteratively position the first audio and the second audio, iteratively determine the audible flight time, and form the ensemble boundary volume. In each iteration, the audio playback server is further configured to adjust the spatial distance to determine the maximum spatial distance until the difference between the audible times of flight exceeds an audible time of flight threshold, as described in item 35. (Item 38) The real-world venue according to item 35, wherein the audio playback server is further configured to move the second audio sound from the second virtual loudspeaker to a third virtual loudspeaker that is less than the maximum spatial distance from the first audio sound during the composite audio program. (Item 39) The movement of the second audio sound from the second virtual loudspeaker to the third virtual loudspeaker is a real-world venue as described in item 38, comprising the snap spatial movement of the second audio sound. (Item 40) The aforementioned audio playback server To isolate the attack transient and decay transient of the first audio from the first audio, Assigning the attack transient of the first audio sound to the first virtual loudspeaker, and assigning the decay transient of the first audio sound to a third virtual loudspeaker from among the plurality of virtual loudspeakers that is less than the maximum spatial distance from the first audio sound, During the composite audio program, the attack transient of the first audio sound is moved from the first virtual loudspeaker to a fourth virtual loudspeaker that is less than the maximum spatial distance from the first audio sound, or During the composite audio program, the decay transient of the first audio sound is moved from the third virtual loudspeaker to the fourth virtual loudspeaker. A real-world venue as described in item 35, further configured to perform the following actions. [Brief explanation of the drawing]
[0007] The accompanying drawings incorporated herein and forming part of the specification illustrate and describe the disclosure, and further illustrate its principles, enabling those skilled in the art to manufacture and use it.
[0008] [Figure 1] Figure 1 illustrates a high-level graphical representation of an exemplary audio system that may be utilized by an exemplary real-world venue according to some exemplary embodiments of the present disclosure.
[0009] [Figure 2] Figure 2 illustrates the operation of an exemplary audio demixing tool within an exemplary audio system for decomposing a composite audio program, according to some exemplary embodiments of the present disclosure.
[0010] [Figure 3] Figure 3 illustrates the operation of an exemplary audio demixing tool for analyzing a composite audio program according to some exemplary embodiments of the present disclosure.
[0011] [Figure 4] Figure 4 illustrates a flowchart of an exemplary audio demixing tool according to several exemplary embodiments of the present disclosure.
[0012] [Figure 5] Figure 5 illustrates in diagram form an exemplary ensemble boundary volume that may be generated by an exemplary audio remixing tool in an exemplary audio system according to some exemplary embodiments of the present disclosure.
[0013] [Figure 6] Figure 6 illustrates an exemplary virtual venue that can be accessed by an exemplary audio remixing tool in some exemplary embodiments of the present disclosure.
[0014] [Figure 7A]Figures 7A–7F illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation to play a composite audio program on a real-world loudspeaker in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 7B] Figures 7A–7F illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation to play a composite audio program on a real-world loudspeaker in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 7C] Figures 7A–7F illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation to play a composite audio program on a real-world loudspeaker in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 7D] Figures 7A–7F illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation to play a composite audio program on a real-world loudspeaker in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 7E] Figures 7A–7F illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation to play a composite audio program on a real-world loudspeaker in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 7F] Figures 7A–7F illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation to play a composite audio program on a real-world loudspeaker in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure.
[0015] [Figure 8A]Figures 8A–8B illustrate the operation of an exemplary audio remixing tool in constructing an exemplary dynamic audio presentation to play a composite audio program on a real-world loudspeaker in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. [Figure 8B] Figures 8A–8B illustrate the operation of an exemplary audio remixing tool in constructing an exemplary dynamic audio presentation to play a composite audio program on a real-world loudspeaker in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure.
[0016] [Figure 9] Figure 9 illustrates, in diagrammatic form, the high-level graphical mapping of an exemplary virtual venue to an exemplary real-world venue in some exemplary embodiments of the present disclosure.
[0017] [Figure 10] Figure 10 illustrates a flowchart of an exemplary audio remixing tool according to some exemplary embodiments of the present disclosure.
[0018] [Figure 11] Figure 11 illustrates a simplified block diagram of a computing device that may be used to implement an electronic device in an exemplary real-world venue according to some embodiments of the present disclosure.
[0019] In accompanying drawings, similar reference numbers indicate the same or functionally similar elements. Additionally, the leftmost digit of the reference number identifies the drawing in which the reference number first appears. [Modes for carrying out the invention]
[0020] (An exemplary audio system for use in an exemplary real-world venue) Figure 1 illustrates a high-level graphical representation of an exemplary audio system that may be utilized by an exemplary real-world venue in some exemplary embodiments of the present disclosure. In the exemplary embodiments illustrated in Figure 1, the audio playback system 100 can play a composite audio program associated with an event being held at a real-world venue 102. For example, the real-world venue 102 may represent a music real-world venue, e.g., a music theater, a music club, and / or a concert hall; a sports real-world venue, e.g., an arena, a convention center, and / or a stadium; and / or any other suitable real-world venue that would be obvious to those skilled in the art without departing from the spirit and scope of the present disclosure. In another embodiment, the event may represent a music event, a theatrical event, a sporting event, a video, and / or any other suitable event that would be obvious to those skilled in the art without departing from the spirit and scope of the present disclosure. In some embodiments, the audio playback system 100 may have access to a composite audio program. As described herein, the audio playback system 100 can run an audio demixing tool to seamlessly decompose a composite audio program into multiple audio sounds that can be collectively played back by the real-world venue 102. Also as described herein, the audio playback system 100 can run an audio remixing tool to intelligently construct an audio presentation from these multiple audio sounds and play back the composite audio program within the real-world venue 102. In some embodiments, the audio playback system 100 may include an audio playback server 104 to perform audio demixing and / or audio remixing.
[0021] In the exemplary embodiment illustrated in Figure 1, the audio playback server 104, whose exemplary embodiment is described in more detail below, can run an audio demixing tool 152 to seamlessly decompose a composite audio program 150 into a plurality of audio sounds that can be collectively played by a real-world venue 102. Alternatively, or in addition, the audio playback server 104 can run an audio remixing tool 154 to intelligently construct an audio presentation from the plurality of audio sounds, play these audio sounds within the real-world venue 102, and play the composite audio program 150 within the real-world venue 102. The audio demixing tool 152 and / or audio remixing tool 154, described in more detail below, can represent one or more software tools that can be executed by one or more electrical, mechanical, and / or electromechanical devices, which will be obvious to those skilled in the art without departing from the spirit and scope of this disclosure. Those skilled in the art will recognize that embodiments of the disclosure described herein can be implemented in hardware, firmware, software, or any combination thereof without departing from this disclosure. Furthermore, those skilled in the art will recognize that firmware, software, routines, instructions, or equivalents may be described herein as performing certain actions. However, it should be understood that such descriptions are for convenience only, and that such actions are actually brought about by one or more electrical, mechanical, and / or electromechanical devices executing the firmware, software, routines, instructions, or equivalents. Alternatively, or in addition, those skilled in the art will recognize that embodiments of the disclosure described herein may also be implemented without departing from the disclosure as instructions stored on a machine-readable medium that can be read and executed by one or more processors. The machine-readable medium may include, for example, any mechanism for storing in a machine-readable form, such as a computing device.For example, machine-readable media may include read-only memory (ROM), random-access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, and equivalents.
[0022] As shown in Figure 1, the audio playback server 104 can run an audio demixing tool 152 to seamlessly decompose the composite audio program 150 into audio sounds 156.1-156.n. In some embodiments, audio sounds 156.1-156.n may include one or more audio channels, for example, a stereo audio channel that may include two audio channels, i.e., a left mono audio channel and a right audio channel feed, or a 5.1 surround audio channel that may include, among other things, a left mono audio channel, a center mono audio channel, a right mono audio channel, a left mono audio channel, and / or a right mono audio channel. In some embodiments, audio sounds 156.1-156.n may include sounds generated by an audio source such as an electronic, mechanical, and / or electromechanical device, which will be obvious to those skilled in the art without departing from the spirit and scope of this disclosure. For example, these electronic, mechanical, and / or electromechanical devices may include, for example, simple instruments such as snare drums, and / or more complex instrument sets such as a standard drum kit having a snare drum, bass drum, one or more tom-toms, one or more cymbals, and / or one or more hi-hat cymbals. In some embodiments, simple instruments may include, for example, percussion instruments, wind instruments, string instruments, and / or electronic instruments. Also, in some embodiments, instrument sets may include instruments from the same classification of instruments, such as percussion instruments, wind instruments, string instruments, and / or electronic instruments, and / or instruments from different classifications. Alternatively, or in addition, audio sounds 156.1–156.n may include natural audio sounds produced by non-human and / or human organisms, such as musical audio sounds produced using a human voice, often referred to as vocals.These natural audio sounds may also include natural non-living sources such as water and / or lightning, to name a few. In some embodiments, the audio demixing tool 152 can analyze the audio sounds 156.1–156.n and identify one or more audio sources that produced these audio sounds, including, among other things, electronic, mechanical, and / or electromechanical devices, one or more non-human organisms, and / or one or more human organisms.
[0023] As described herein, the audio demixing tool 152 can decompose the composite audio program 150 into audio sounds 156.1–156.n. In some embodiments, the audio demixing tool 152 can analyze the composite audio program 150 and identify the audio sounds 156.1–156.n present within it. As part of this identification, the audio demixing tool 152 can iteratively explore the composite audio program 150 with respect to the audio sounds 156.1–156.n. For example, the audio demixing tool 152 can explore the composite audio program 150 with respect to audio sounds produced by a snare drum, audio sounds produced by a bass guitar, and so on. In some embodiments, the audio demixing tool 152 can iteratively explore the composite audio program 150 starting from the most dominant or dominant audio sound among the audio sounds 156.1–156.n. After identifying audio sounds 156.1–156.n, the audio demixing tool 152 can decompose the composite audio program 150 into audio sounds 156.1–156.n. As part of this decomposition, the audio demixing tool 152 can iteratively isolate audio sounds 156.1–156.n from the composite audio program 150 and provide the corresponding audio sounds from audio sounds 156.1–156.n. After this decomposition, the audio demixing tool 152 can store the corresponding audio sounds for reading by an audio remixing tool 154 as described herein.
[0024] As illustrated in Figure 1, the audio remixing tool 154 can intelligently construct an audio presentation from audio sounds 156.1-156.n and play a composite audio program 150 within the real-world venue 102. In some embodiments, the audio remixing tool 154 can construct the audio presentation in advance, for example, offline prior to an event, and / or during an event, for example, in real time or near real time and concurrently with the event. In some embodiments, once the audio presentation is played back by the audio playback server 104, the real-world venue 102 can be configured to play the audio sounds 156.1-156.n on real-world loudspeakers within the real-world venue 102. In the exemplary embodiment illustrated in Figure 1, the audio remixing tool 154 can assign the audio sounds 156.1-156.n to real-world loudspeakers within the real-world venue 102 and construct an audio presentation for playing a composite audio program 150 within the real-world venue 102. In some embodiments, the audio presentation can represent a static audio presentation, where the assignment of audio voices 156.1-156.n to real-world loudspeakers in a real-world venue 102 remains fixed or static throughout the composite audio program 150, and / or a dynamic audio presentation, where the assignment of audio voices 156.1-156.n to real-world loudspeakers in a real-world venue 102 moves, changes, or switches dynamically during the composite audio program. In some embodiments, the audio remixing tool 154 may take into account one or more characteristics, parameters, and / or attributes of the audio voices 156.1-156.n when constructing the audio presentation.In some embodiments, these characteristics, parameters, and / or attributes can be utilized by the audio remixing tool 154 to provide a more realistic and aesthetically pleasing reproduction of the composite audio program 150, while simultaneously mitigating the effects of comb filtering and / or audio latency as described herein within the real-world venue 102. Therefore, mitigating the effects of comb filtering and / or audio latency as described herein within the real-world venue 102 may be important to ensure a seamless and immersive experience.
[0025] After assigning audio signals 156.1-156.n, the audio remixing tool 154 can identify audio control signals 158.1-158.i to configure the audio equipment in the real-world venue 102 to play the audio signals 156.1-156.n through real-world loudspeakers in the real-world venue 102. In some embodiments, the real-world venue 102 may include audio equipment such as amplifiers, crossovers, equalizers, and / or mixers to route and / or signal-modulate the audio signals 156.1-156.n for playback within the real-world venue 102. In these embodiments, the audio remixing tool 154 can generate audio control signals 158.1-158.i and configure the audio equipment in the real-world venue 102 to play the audio signals 156.1-156.n on real-world loudspeakers in the real-world venue 102. In these embodiments, audio control signals 158.1-158.i can cause audio equipment in the real-world venue 102 to route and / or signal-modulate audio sounds 156.1-156.n for playback through real-world loudspeakers in the real-world venue 102, as defined in the audio presentation.
[0026] In the exemplary embodiment illustrated in Figure 1, the real-world venue 102 can represent a three-dimensional structure, for example, a hemispherical structure, also referred to as a hemispherical dome. In some embodiments, the real-world venue 102 can include one or more visual displays, often referred to as a three-dimensional media surface, which are diffused across the interior of the real-world venue 102. In these embodiments, one or more visual displays can include a series of rows and columns of image elements, also referred to as pixels, in three dimensions, which form the three-dimensional media surface for projecting an image or a series of images onto the three-dimensional media surface, often referred to as a video, which may be associated with an event, for example. In these embodiments, pixels can be implemented using, to name a few, one or more light-emitting diode (LED) displays, one or more organic light-emitting diode (OLED) displays, and / or one or more quantum dot (QD) displays. For example, the three-dimensional media surface could include a 19,000 × 13,500 LED visual display that surrounds the interior of the real-world venue 102 and forms a visual display of approximately 160,000 square feet.
[0027] Furthermore, as illustrated in Figure 1, the real-world venue 102 can play audio sounds 156.1-156.n within the real-world venue 102 as defined in the audio presentation, and can play a composite audio program 150 within the real-world venue 102. As illustrated in Figure 1, the real-world venue 102 may include real-world loudspeakers 106.1-106.i for playing audio sounds 156.1-156.n within the real-world venue 102. Alternatively, or in addition, the real-world venue 102 may include audio equipment as described herein for routing audio sounds 156.1-156.n to the real-world loudspeakers 106.1-106.i and / or for signal conditioning of audio sounds 156.1-156.n. In some embodiments, the audio remixing tool 154 can reproduce audio sounds 156.1-156.n through real-world loudspeakers 106.1-106.i as defined in the audio presentation, in a manner substantially similar to that described herein. In some embodiments, the real-world loudspeakers 106.1-106.i may include a proscenium array real-world loudspeaker system installed in or near the proscenium of a real-world venue 102, one or more effect extension array real-world loudspeaker systems installed in or near the proscenium array real-world loudspeaker system, and / or one or more environment array real-world loudspeaker systems installed throughout the entire real-world venue 102. In some embodiments, a proscenium array real-world loudspeaker system, one or more effect-enhanced array real-world loudspeaker systems, and / or one or more environment array real-world loudspeaker systems may include one or more real-world loudspeakers, which may include, to some extent, one or more super tweeters, one or more tweeters, one or more midrange speakers, one or more woofers, one or more subwoofers, and / or one or more full-range speakers. (An exemplary audio demixing tool for disassembling complex audio programs)
[0028] Figure 2 illustrates the operation of an exemplary audio demixing tool in an exemplary audio system for demixing a composite audio program according to some exemplary embodiments of the present disclosure. In the exemplary embodiments illustrated in Figure 2, the audio demixing tool 200 can access a composite audio program, such as the composite audio program 150. As described herein, the audio demixing tool 200 can seamlessly demix the composite audio program into a plurality of audio sounds that can be collectively played by a real-world venue. In some embodiments, the audio demixing tool 200 may represent one or more software tools that can be executed by one or more electrical, mechanical, and / or electromechanical devices, such as the audio playback server 104, without departing from the spirit and scope of the present disclosure, in order to seamlessly demix the composite audio program into a plurality of audio sounds. The audio demixing tool 200 as described herein may represent an exemplary embodiment of the audio demixing tool 152.
[0029] In the exemplary embodiment illustrated in Figure 2, the audio demixing tool 200 can decompose the composite audio program 202 into audio voices 208.1-208.m through a process called music source separation. In some embodiments, the audio voices 208.1-208.m may include one or more audio channels, for example, a stereo audio channel that may include two audio channels, i.e., a left mono audio channel and a right audio channel feed, or a 5.1 surround audio channel that may include, among other things, a left mono audio channel, a center mono audio channel, a right mono audio channel, a left mono audio channel, and / or a right mono audio channel. As part of this music source separation, the audio demixing tool 200 can analyze the composite audio program 202 and identify from the composite audio program 202 the matching audio sounds 208.1-208.m corresponding to the sample audio sounds 204.1-204.r in the electronic library of audio sounds 206. The sample audio sounds 204.1-204.r may include, among other things, audio sounds produced by one or more audio sources such as one or more electronic, mechanical, and / or electromechanical devices, one or more non-human organisms, and / or one or more human organisms, as described herein. For example, sample audio sounds 204.1–204.r may include, among other things, audio sample 204.1 produced by a microphone, audio sample 204.2 produced by an acoustic guitar, electric guitar, and bass guitar, audio sample 204.a produced by a drum kit, audio sample 204.a+1 produced by the bass drum from the drum kit, and audio sample 204.r produced by the snare drum from the drum kit.
[0030] In some embodiments, the audio demixing tool 200 can iteratively search the composite audio program 202 for the presence of sample audio sounds 204.1–204.r. For example, the audio demixing tool 200 can iteratively search the composite audio program 202 for the presence of, among other things, audio sample 204.1 produced by a microphone, audio sample 204.2 produced by an acoustic guitar, an electric guitar, and a bass guitar, audio sample 204.a produced by a drum kit, audio sample 204.a+1 produced by a bass drum from a drum kit, and audio sample 204.r produced by a snare drum from a drum kit. In the exemplary embodiment illustrated in Figure 2, the audio demixing tool 200 can execute a pattern recognition algorithm to iteratively search the composite audio program 202 for sample audio sounds 204.1–204.r. In some embodiments, the pattern recognition algorithm may include a template matching algorithm that matches sample audio speech 204.1-204.r to a composite audio program 202, and / or a structure / syntax matching algorithm or a statistical matching algorithm, respectively, with semi-supervised and supervised machine learning and subsequent matching of the sample audio speech 204.1-204.r to the composite audio program 202.
[0031] As part of the music source separation, the audio demixing tool 200 can decompose the composite audio program 202 into audio sounds 208.1-208.m. After identifying the sample audio sounds 204.1-204.r, the audio demixing tool 200 can iteratively isolate each of the sample audio sounds 204.1-204.r present in the composite audio program 202 to provide audio sounds 208.1-208.m. In some embodiments, the audio demixing tool 200 can iteratively subtract each of the sample audio sounds 204.1-204.r present in the composite audio program 202 to isolate each of the sample audio sounds 204.1-204.r present in the composite audio program 202. From the above embodiment, the audio demixing tool 200 can, in particular, identify that audio sample 204.1 generated by a microphone, audio sample 204.2 generated by an acoustic guitar, electric guitar, and bass guitar, and audio sample 204.a generated by a drum kit are present in the composite audio program 202. In this embodiment, the audio demixing tool 200 can, in particular, isolate the audio sound generated by the microphone from the composite audio program 202 and provide a first audio sound from audio sounds 208.1-208.b, isolate the audio sound generated by the acoustic guitar, electric guitar, and bass guitar and provide a second audio sound from audio sounds 208.1-208.b, and / or isolate the audio sound generated by the drum kit from the composite audio program 202 and provide a bth audio sound from audio sounds 208.1-208.b.
[0032] In some embodiments, the audio demixing tool 200 can hierarchically decompose the composite audio program 202 into audio sounds 208.1-208.m. In these embodiments, the audio demixing tool 200 can analyze the composite audio program 202 and identify sample audio sounds corresponding to sets of instruments from among the sample audio sounds 204.1-204.r, such as audio sample 204.a generated by a drum kit, which are present within the composite audio program 202. After identifying these sample audio sounds, the audio demixing tool 200 can iteratively isolate each of these sample audio sounds present within the composite audio program 202 and provide audio sounds corresponding to sets of instruments from among the audio sounds 208.1-208.m, such as audio sound 208.b, which are present, for example. Subsequently, the audio demixing tool 200 can again decompose audio 208.b into audio 208.c-208.m and provide audio 208.1-208.m. The audio demixing tool 200 can analyze audio 208.b and identify sample audio 204.1-204.r present within audio 208.b. After identifying the sample audio 204.1-204.r present within audio 208.b, the audio demixing tool 200 can iteratively isolate each of the sample audio 204.1-204.r present within audio 208.b from audio 208.b and provide audio 208.c-208.m from audio 208.1-208.m.
[0033] In some embodiments, the audio demixing tool 200 can analyze audio sounds 208.1–208.m and identify one or more audio sources that generated these audio sounds, such as, among others, one or more electronic, mechanical, and / or electromechanical devices, one or more non-human beings, and / or one or more human beings, as described herein. In these embodiments, the audio demixing tool 200 can analyze one or more of the audio sounds 208.1–208.m to identify, for example, whether a simple instrument such as a snare drum generated these audio sounds. Alternatively, or in addition, one or more of the audio sounds 208.1–208.m can be analyzed to identify, for example, whether a more complex collection of instruments such as a standard drum kit having a snare drum, bass drum, one or more tom-toms, one or more cymbals, and / or one or more hi-hat cymbals generated these audio sounds. From the above embodiment, the audio demixing tool 200 can analyze a first audio sound from among audio sounds 208.1-208.m and identify that a microphone generated this audio sound; analyze a second audio sound from among audio sounds 208.1-208.m and identify that an acoustic guitar, electric guitar, and bass guitar generated this audio sound; and analyze a b-th audio sound from among audio sounds 208.1-208.m and identify that a drum kit generated this audio sound. (An exemplary audio demixing tool for analyzing audio from a composite audio program)
[0034] Figure 3 illustrates the operation of an exemplary audio demixing tool for analyzing a composite audio program according to some exemplary embodiments of the present disclosure. In the exemplary embodiments illustrated in Figure 3, the audio demixing tool 300 can access multiple audio sounds of a composite audio program, such as the composite audio program 150. As described herein, the audio demixing tool 300 can analyze the multiple audio sounds and identify one or more characteristics, parameters, and / or attributes of these audio sounds. As described herein, one or more characteristics, parameters, and / or attributes can be used by an audio remixing tool, such as the audio remixing tool 154, to construct an audio presentation for playing the composite audio program in a real-world venue. In some embodiments, the audio remixing tool can utilize one or more characteristics, parameters, and / or attributes to provide a more realistic and aesthetically pleasing playback of the composite audio program while mitigating the effects of comb filtering and / or audio latency as described herein in a real-world venue. In these embodiments, mitigating the effects of comb filtering and / or audio latency as described herein within a real-world venue may be important to ensure a seamless and immersive experience.
[0035] As illustrated in Figure 3, the audio demixing tool 300 can access audio sounds 302.1-302.m decomposed from a composite audio program in a manner substantially similar to that described herein. In some embodiments, audio sounds 302.1-302.m can represent, among other things, sounds produced by electronic, mechanical, and / or electromechanical devices, one or more non-human beings, and / or one or more human beings, as described herein. After accessing the audio sounds 302.1-302.m, the audio demixing tool 300 can analyze the audio sounds 302.1-302.m and identify one or more characteristics, parameters, and / or attributes 306.1-306.m corresponding to these audio sounds. As illustrated in Figure 3, the audio demixing tool 300 may include an audio voice analysis software engine 304 to process audio voices 302.1-302.m and identify one or more characteristics, parameters, and / or attributes 306.1-306.m. For example, the audio voice analysis software engine 304 may, among other things, process audio voice 302.1 and identify one or more characteristics, parameters, and / or attributes 306.1 corresponding to audio voice 302.1, and process audio voice 302.2 and identify one or more characteristics, parameters, and / or attributes 306.2 corresponding to audio voice 302.2. In some embodiments, one or more characteristics, parameters, and / or attributes 306.1-306.m may, among other things, include the pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity of the audio voices 302.1-302.m. In these embodiments, the audio speech analysis software engine 304 can compare the pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity between audio speeches 302.1–302.m and identify one or more characteristics, parameters, and / or attributes 306.1–306.m.In some embodiments, the audio speech analysis software engine 304 may store one or more characteristics, parameters, and / or attributes 306.1–306.m as an organized collection of data, often referred to as a database. The database may include one or more data tables having data values such as alphanumeric strings, integers, decimals, floating-point numbers, dates, times, binary values, Boolean values, and / or enumerations, to name a few. The database may be a column-oriented database, a relational database, a keystore database, a graph database, and / or a document store, to name a few. In these embodiments, an audio remixing tool may access the database, read one or more characteristics, parameters, and / or attributes 306.1–306.m, and construct an audio presentation for playing a composite audio program in a real-world venue as described herein.
[0036] In some embodiments, one or more characteristics, parameters, and / or attributes 306.1-306.m may indicate the spatial positioning between one or more audio sources that generated the audio sounds 302.1-302.m, the audio transients of the audio sounds 302.1-302.m, the spatial movement of one or more audio sources that generated the audio sounds 302.1-302.m, the timing relationships between the audio sounds 302.1-302.m, the audio effects within the audio sounds 302.1-302.m, and / or any other suitable characteristics, parameters, and / or attributes of or between the audio sounds 302.1-302.m that would be recognized by those skilled in the art without departing from the spirit and scope of this disclosure.
[0037] The spatial positioning between one or more audio sources that produced audio sounds 302.1–302.m indicates the relative positioning between two or more audio sources that produced two or more of the audio sounds 302.1–302.m relative to each other. In these embodiments, the audio sound analysis software engine 304 can perform a three-dimensional audio sound localization technique between two or more of the audio sounds 302.1–302.m to estimate the spatial positioning between two or more audio sources that produced these audio sounds. In these embodiments, the audio sound analysis software engine 304 can estimate the interaural time difference (ITD) and / or interaural intensity difference (IID) between two or more of the audio sounds 302.1–302.m to estimate the relative positioning between two or more audio sources.
[0038] The audio transients of audio voices 302.1–302.m represent the envelope of audio voices 302.1–302.m over time. In some embodiments, these envelopes may represent the attack transient, decay transient, sustain transient, and / or release transient of audio voices 302.1–302.m. In these embodiments, the attack transient represents a first duration until the audio voices 302.1–302.m reach their maximum amplitude, the decay transient represents a second duration until the audio voices 302.1–302.m decrease from their maximum amplitude to their steady-state amplitude, the sustain transient represents a third duration until the audio voices 302.1–302.m are at their steady-state amplitude, and the release transient represents a fourth duration until the audio voices 302.1–302.m decrease from their steady-state amplitude to their minimum amplitude.
[0039] The spatial movement of one or more audio sources that produced audio speeches 302.1-302.m represents the relative movement between two or more audio sources that produced two or more of the audio speeches 302.1-302.m relative to each other. In the exemplary embodiment illustrated in Figure 3, two or more audio sources may change location or move around during the composite audio program. In some embodiments, the audio speech analysis software engine 304 can analyze the audio speeches 302.1-302.m to determine whether two or more audio sources are moving. For example, the audio demixing tool 300 can compare the amplitude and / or phase of two or more of the audio speeches 302.1-302.m corresponding to two or more audio sources to estimate whether two or more audio sources are moving relative to each other.
[0040] The timing relationships between audio sounds 302.1–302.m represent the relative timing relationships between two or more audio sources that produced two or more of the audio sounds 302.1–302.m relative to each other. In some embodiments, the timing relationships may be referred to as beat relationships, bar relationships, and / or tick relationships between two or more audio sources. These relative timing relationships may include time signatures such as simple time signatures, compound time signatures, pulsating time signatures, common time signatures, complex time signatures, mixed time signatures, additive time signatures, irrational time signatures, and / or equivalents. In some embodiments, the audio analysis software engine 304 may run a timing relationship algorithm to compare two or more of the audio sounds 302.1–302.m relative to each other and identify the relative timing relationships between two or more audio sources. As part of the timing relationship algorithm, the audio analysis software engine 304 may classify two or more audio sources. In some embodiments, the audio analysis software engine 304 can classify each of two or more audio sources according to a general type, e.g., percussion instruments, wind instruments, string instruments, and / or electronic instruments. The audio analysis software engine 304 can then identify the relative timing relationships between the two or more audio sources according to their general types. For example, a percussion instrument from among two or more audio sources may have the same timing relationship as another percussion instrument from among two or more audio sources, while a percussion instrument may have a different timing relationship with a wind instrument from among two or more audio sources. Alternatively, or in addition, the audio analysis software engine 304 can classify each of two or more audio sources according to a specific type, e.g., snare drum, bass drum, tom-tom, cymbal. The audio analysis software engine 304 can then identify the relative timing relationships between the two or more audio sources according to their specific types.For example, a snare drum from two or more audio sources may have the same timing relationship as another snare drum from two or more audio sources, while a snare drum may have a different timing relationship than a cello from two or more audio sources.
[0041] The audio effects within audio 302.1-302.m represent specific audio effects that may be applied by one or more audio sources that generated audio 302.1-302.m in order to modify audio 302.1-302.m. In the exemplary embodiment illustrated in Figure 3, the audio analysis software engine 304 can analyze audio 302.1-302.m and identify the audio effects within it. In some embodiments, one or more audio effects may include, to name a few, one or more echo, flanger, phaser, chorus, equalization, filtering, overdrive, pitch shift, time stretch, resonator, voice effect, synthesizer, modulation, compression, and / or equivalents. In some embodiments, the audio analysis software engine 304 can isolate the audio generated by the audio source, referred to as the parent audio, from the audio effects within audio 302.1-302.m, referred to as the child audio. (Example behavior of an exemplary audio demixing tool)
[0042] Figure 4 illustrates a flowchart of an exemplary audio demixing tool according to several exemplary embodiments of the present disclosure. The present disclosure is not limited to this description of operation. Rather, it will be apparent to those skilled in the art that other operation control flows are also within the scope and spirit of the present disclosure. The following discussion describes an operation control flow 400 for seamlessly decomposing a composite audio program, such as composite audio program 150, and / or identifying the characteristics, parameters, and / or attributes corresponding to the audio sounds of the composite audio program. The operation control flow 400 can be performed, for example, by an audio demixing tool 152. In some embodiments, the operation control flow 400 can be performed by one or more computing devices, such as an audio playback server 104.
[0043] In operation 402, the operation control flow 400 decomposes the composite audio program in a manner substantially similar to that described herein, providing audio sounds such as audio sounds 156.1-156.n, audio sounds 208.1-208.m, and / or audio sounds 302.1-302.m.
[0044] In operation 404, the operation control flow 400 can analyze one or more of the audio sounds from operation 402 and identify one or more characteristics, parameters, and / or attributes of these audio sounds in a manner substantially similar to that described herein. These characteristics, parameters, and / or attributes can also be utilized by an audio remixing tool, such as audio remixing tool 154, to construct an audio presentation for playing a composite audio program in a real-world venue, as described herein. (Example of the operation of an exemplary audio remixing tool)
[0045] Before describing exemplary audio remixing tools that may be implemented in the exemplary real-world venue described herein, an audio ensemble will be described in general terms. As described herein, a composite audio program may include multiple audio sounds produced by multiple audio sources, such as percussion instruments, wind instruments, string instruments, and / or electronic instruments, to name a few. As described herein, these audio sources can be logically grouped together to form an audio ensemble. In some embodiments, audio sources from among multiple audio sources having similar characteristics, parameters, and / or attributes can be logically grouped together to form an audio ensemble. In these embodiments, to name a few, audio sources from among multiple audio sources having similar pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity can be logically grouped together to form an audio ensemble. For example, a snare drum, a bass drum, one or more tom-toms, one or more cymbals, and / or one or more hi-hat cymbals having similar pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity can be logically grouped together to form an audio ensemble associated with a drum kit.
[0046] Figure 5 illustrates exemplary ensemble boundary volumes that may be generated by an exemplary audio remixing tool in an exemplary audio system according to some exemplary embodiments of the present disclosure. In the exemplary embodiments illustrated in Figure 5, multiple audio sources having similar characteristics, parameters, and / or attributes can be logically grouped together to form an audio ensemble. As described herein, these audio ensembles can be associated with ensemble boundary volumes. These ensemble boundary volumes can define spatial distances between audio sources in the audio ensemble to provide a more realistic and aesthetically pleasing reproduction of a composite audio program, such as composite audio program 150, while simultaneously mitigating the effects of comb filtering and / or audio latency as described herein in a real-world venue, such as real-world venue 102. Thus, mitigating the effects of comb filtering and / or audio latency as described herein in a real-world venue may be important to ensure a seamless and immersive experience. The discussion in Figure 5 that follows illustrates an exemplary ensemble boundary volume relating to a drum kit, which may include, to some extent, a snare drum, a bass drum, one or more tom-toms, one or more cymbals, and / or one or more hi-hat cymbals. Those skilled in the art will recognize that other ensemble boundary volumes relating to drum kits and / or other audio sources may also be implemented in a substantially similar manner to those described herein without departing from the spirit and scope of this disclosure.
[0047] As illustrated in Figure 5, audio sources 500.1–500.x having similar characteristics, parameters, and / or attributes from among multiple audio sources of a composite audio program, for example, a snare drum, bass drum, one or more tom-toms, one or more cymbals, and / or one or more hi-hat cymbals from a drum kit, can be logically grouped together to form an audio ensemble 502. In the exemplary embodiment illustrated in Figure 5, the audio ensemble 502 having audio sources 500.1–500.x can be associated with an ensemble boundary volume 504. The ensemble boundary volume 504 is illustrated in Figure 5 as a bounding box in three-dimensional space, but this is for illustrative purposes only and is not intended to be limiting. Those skilled in the art will recognize that the ensemble boundary volume 504 may be any suitable three-dimensional volume in three-dimensional space, such as a bounding capsule, bounding cylinder, bounding ellipsoid, bounding sphere, bounding slab, and / or bounding triangle, without departing from the spirit and scope of this disclosure. Furthermore, although the ensemble boundary volume 504 is illustrated in Figure 5 as being in three-dimensional space, those skilled in the art will recognize that the ensemble boundary volume 504 can similarly be implemented in two-dimensional space as any suitable two-dimensional shape, e.g., a circle, triangle, quadrilateral, and / or polygon, to form the ensemble boundary area without departing from the spirit and scope of the disclosure. In some embodiments, the audio sources 500.1–500.x can be characterized as diffuse audio sources, transient audio sources, and / or any combination of diffuse and transient audio sources. In these embodiments, the diffuse audio source represents an audio source that produces sound that develops gradually over a relatively long duration, while the transient audio source produces sound that develops abruptly over a relatively short duration.
[0048] Generally, the ensemble boundary volume 504 defines the spatial distance between audio sources 500.1 - 500.x in three-dimensional space. In the exemplary embodiment illustrated in FIG. 5, the ensemble boundary volume 504 defines the maximum spatial distance between audio sources 500.1 - 500.x in three-dimensional space. In the exemplary embodiment illustrated in FIG. 5, the maximum spatial distance between audio sources 500.1 - 500.x represents the spatial distance between audio sources 500.1 - 500.x that is less than or equal to the audible flight time threshold and has a difference in audible flight time equal to the audible flight time threshold. In some embodiments, the audible flight time threshold is such that the difference between the first audible flight time T A,B and the second audible flight time T B,1 can be selectively chosen to be less than the temporal resolution of human hearing, e.g., between about 20 (20) milliseconds to about 36 (36) milliseconds. In these embodiments, the effects of comb filtering and / or audio latency are typically not noticed when the difference between the first audible flight time T A,1 and the second audible flight time T B,1 is less than the temporal resolution of human hearing as described herein.
[0049] As illustrated in FIG. 5, the audio demixing tool determines the distance D A , Y A , Z A ) and the second three-dimensional coordinates (X B , Y B , Z B ) within the ensemble boundary volume 504 associated with the audio source 500.x. Thereafter, the audio demixing tool determines the first audible flight time T for the first audio sound generated by the audio source 500.1 to propagate from the first three-dimensional coordinates (X A , Y A , Z A ) to the first three-dimensional coordinates (x A , y A , z A ) in three-dimensional space.
[0049] <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> <000********> [[ID=228A,1 Then, the second audio sound generated by audio source 500.x is in the second 3D coordinate (X) in 3D space. B ,Y B ,Z B ) from the first 3D coordinate (x A ,y A ,z A The second audible flight time T until it propagates to ) B,1 It is possible to estimate the first audible time of flight T. In some embodiments, the audio demixing tool estimates the first audible time of flight T. A,1 It can be estimated as follows: [ka] Second audible flight time T B,1 It can be estimated as follows: [ka] T1 and T2 are the first audible flight time T, respectively. A,1 and the second audible flight time T B,1 D1 and D2 represent the first three-dimensional coordinate (X A ,Y A ,Z A ) and the first 3D coordinate (x A ,y A ,z A ) and the second 3D coordinate (X B ,Y B ,Z B ) and the first 3D coordinate (x A ,y A ,z A This represents the distances D1 and D2 between ) and v sound This represents the speed of sound. Typically, the speed of sound in air is approximately 343 meters / second at 20 degrees Celsius, and this can vary depending on the temperature. Subsequently, the audio demixing tool uses the first audible time of flight T A,1 and the second audible flight time T B,1The difference between the two can be estimated. In some embodiments, the audio demixing tool uses the first audible time of flight T A,1 and the second audible flight time T B,1 The difference between the first audible time of flight T can be compared to the audible time of flight threshold. In these embodiments, the audio demixing tool uses the first audible time of flight T A,1 and the second audible flight time T B,1 When the difference between audio source 500.1 and audio source 500.x is less than or equal to the audible time-of-flight threshold, the distance D A,B This makes it possible to separate them spatially in three-dimensional space. In these embodiments, the distance D A,B This is the first audible flight time T A,1 and the second audible flight time T B,1 When the difference between and is less than or equal to the audible time-of-flight threshold, it can be characterized as being contained within the ensemble boundary volume 504. Otherwise, in some embodiments, the audio demixing tool is contained within the first audible time-of-flight T A,1 and the second audible flight time T B,1 When the difference between audio source 500.1 and audio source 500.x exceeds the audible time-of-flight threshold, distance D A,B This prevents them from being separated from each other and spatially distanced in three-dimensional space. In these embodiments, distance D A,B This is the first audible flight time T A,1 and the second audible flight time T B,1When the difference between the two is less than or equal to the audible time-of-flight threshold, it can be characterized as being outside the ensemble boundary volume 504. In some embodiments, the audio demixing tool can empirically simulate audio sources 500.1 and 500.x in different three-dimensional coordinates to determine the faces, vertices, and / or surfaces of the ensemble boundary volume 504 in three-dimensional space. In these embodiments, the audio demixing tool can execute a computation algorithm, such as a Monte Carlo algorithm, to empirically simulate audio sources 500.1 and 500.x in different three-dimensional coordinates.
[0050] Furthermore, the audio demixing tool, in a manner substantially similar to that described herein, processes the first audio sound generated by the audio source 500.1 into a first three-dimensional coordinate (X) in three-dimensional space. A ,Y A ,Z A ) to the second 3D coordinate (x B ,y B ,z B The first audible flight time T until it propagates to ) A,2 Then, the second audio sound generated by audio source 500.x is in the second 3D coordinate (X) in 3D space. B ,Y B ,Z B ) to the second 3D coordinate (x B y B ,z B The second audible flight time T until it propagates to ) B,2 It is possible to estimate the first audible flight time T. A,2 and the second audible flight time T B,2 The difference between the first audible flight time T is A,1 and the second audible flight time T B,1 It is approximately equal to the difference between the two. In these embodiments, the first three-dimensional coordinate (x A ,y A ,z AThe first audio sound generated by audio source 500.1 and the second audio sound generated by audio source 500.x in ) are transmitted to the second 3D coordinate (x B y B ,z B The first audio sound produced by audio source 500.1 and the second audio sound produced by audio source 500.x can be characterized as producing sounds substantially similar to those produced by audio source 500.x in the real-world venue. Therefore, the first audio sound produced by audio source 500.1 and the second audio sound produced by audio source 500.x should be substantially similar to each other in most, if not all, locations within the real-world venue. (An exemplary virtual venue that can be accessed by an exemplary audio demixing tool)
[0051] Figure 6 illustrates an exemplary virtual venue that can be accessed by an exemplary audio remixing tool in some exemplary embodiments of the present disclosure. As described herein, an audio remixing tool, such as audio remixing tool 154, can intelligently construct an audio presentation and play a composite audio program, such as composite audio program 150, on a real-world loudspeaker in a real-world venue, such as a real-world venue. As described herein, an audio remixing tool can access a virtual representation of a real-world venue in three-dimensional space, also referred to as a virtual venue 600, which virtually identifies the three-dimensional coordinates of a real-world loudspeaker in three-dimensional space. Although the virtual venue 600 is illustrated in Figure 6 as being in three dimensions within three-dimensional space, those skilled in the art will recognize that the virtual venue 600 may similarly be in two dimensions within two-dimensional space without departing from the spirit and scope of the present disclosure. Furthermore, as described herein, the audio remixing tool can utilize the virtual venue 600 to assign the audio sounds of the composite audio program 150, such as audio sounds 156.1-156.n, to real-world loudspeakers in the real-world venue, thereby constructing an audio presentation for playing the composite audio program within the real-world venue.
[0052] As illustrated in Figure 6, the virtual venue 600 includes virtual loudspeakers 602.1-602.k installed within the three-dimensional space of the virtual venue 600. However, the configuration and arrangement of the virtual loudspeakers 602.1-602.k within the three-dimensional space of the virtual venue 600 as illustrated in Figure 6 are for illustrative purposes only and are not intended to be limiting. Those skilled in the art will recognize that real-world loudspeakers 602.1-602.k may be configured and arranged differently within the three-dimensional space of the virtual venue 600 without departing from the spirit and scope of this disclosure. In some embodiments, the virtual loudspeakers 602.1-602.k each have three-dimensional coordinates (x1, y1, z1)-(x k ,y k ,zk ) can be located in ). In some embodiments, virtual loudspeakers 602.1-602.k can represent virtual representations of real-world loudspeakers in real-world venues such as real-world venue 102, virtual representations of virtual loudspeakers in real-world venues, and / or any combination of real-world loudspeakers or virtual loudspeakers in real-world venues. In these embodiments, virtual representations of real-world loudspeakers may include a proscenium virtual loudspeaker system 604 installed in or near the proscenium of virtual venue 600, an effects-extended virtual array real-world loudspeaker system 608.1-608.l installed in or near the proscenium virtual loudspeaker system 604, and / or an environment virtual array real-world loudspeaker system 610.1-610.m installed throughout virtual venue 600. In these embodiments, the proscenium virtual loudspeaker system 604 may include virtual loudspeakers 606.1–606.z. In some embodiments, the virtual loudspeakers 606.1–606.z, the effect-extended virtual array real-world loudspeaker systems 608.1–608.l, and / or the environment virtual array real-world loudspeaker systems 610.1–610.m may include, to name a few, one or more virtual super tweeters, one or more virtual tweeters, one or more virtual midrange speakers, one or more virtual woofers, one or more virtual subwoofers, and / or one or more virtual full-range speakers. (An example of a static audio presentation that can be constructed using an exemplary audio demixing tool)
[0053] The exemplary static audio presentations described below in Figures 7A–7F represent an audio presentation in which the assignment of audio voices 156.1–156.n of the composite audio program 150 to real-world loudspeakers in a real-world venue such as real-world venue 102 remains fixed or static throughout the composite audio program. On the other hand, the exemplary dynamic audio presentations described herein in Figures 8A–8B represent an audio presentation in which the assignment of audio voices of the composite audio program to real-world loudspeakers in a real-world venue moves, changes, or switches dynamically during the composite audio program.
[0054] Figures 7A–7F illustrate the operation of an exemplary audio remixing tool in constructing an exemplary static audio presentation to play a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. In the exemplary embodiment illustrated in Figure 7A, an audio remixing tool, such as audio remixing tool 154, can intelligently construct a static audio presentation 700 to play a composite audio program on real-world loudspeakers in a real-world venue, such as real-world venue 102. As described herein, the audio remixing tool can utilize a virtual venue 600 and assign the audio sounds of the composite audio program, such as audio sounds 156.1–156.n of the composite audio program 150, to real-world loudspeakers in the real-world venue to construct the static audio presentation 700. Those skilled in the art will recognize that the static audio presentation 700 as illustrated in Figure 7A is for illustrative purposes only and is not intended to be limiting. Those skilled in the art will recognize that, without departing from the spirit and scope of this disclosure, audio remixing tools can intelligently construct other static audio presentations and play other composite audio programs on real-world loudspeakers in real-world venues.
[0055] The discussion of Figures 7A–7F describes exemplary operations that may be utilized by an audio remixing tool to construct a static audio presentation 700 for playing a composite audio program on real-world loudspeakers in a real-world venue. Those skilled in the art will recognize that these operations may be performed independently or in any combination to construct other static audio presentations for playing other composite audio programs on real-world loudspeakers in a real-world venue, without departing from the spirit and scope of this disclosure. In the exemplary embodiment illustrated in Figure 7A, the audio remix can access an electronic library 702 of audio sounds containing audio sounds 704.1–704.m present in the composite audio program. In some embodiments, an audio demixing tool, such as an audio demixing tool 152, can seamlessly demix the composite audio program and identify the audio sounds 704.1–704.m. As illustrated in Figure 7A, audio sounds 704.1–704.m may include, as described herein, audio sounds produced by one or more audio sources such as one or more electronic, mechanical, and / or electromechanical devices, one or more non-human organisms, and / or one or more human organisms. For example, audio sounds 704.1–704.m may include, among other things, audio sound 704.1 produced by a microphone, audio sound 704.2 produced by an acoustic guitar, electric guitar, and bass guitar, audio sound 704.3 produced by a bass drum, and audio sound 704.m produced by a snare drum.
[0056] In the exemplary embodiment illustrated in Figure 7A, the audio remixing tool can assign audio voices 704.1–704.m to virtual loudspeakers 602.1–602.k in the virtual venue 600 and construct a static audio presentation 700 for playing a composite audio program on real-world loudspeakers in a real-world venue. As described herein, the virtual venue 600 may include a proscenium virtual loudspeaker system 604 having virtual loudspeakers 606.1–606.z, an effects-extended virtual array real-world loudspeaker system 608.1–608.l, and / or an environment virtual array real-world loudspeaker system 610.1–610.m. As illustrated in Figure 7A, the audio remixing tool can construct a static audio presentation 700 by assigning audio 704.1 produced by a microphone to virtual loudspeaker 606.4 from the proscenium virtual loudspeaker system 604, audio 704.2 produced by an acoustic guitar, electric guitar, and bass guitar to virtual loudspeaker 606.7 from the proscenium virtual loudspeaker system 604, audio 704.3 produced by a bass drum to virtual loudspeaker 606.2 from the proscenium virtual loudspeaker system 604, and / or audio 704.m produced by a snare drum to virtual loudspeaker 606.1 from the proscenium virtual loudspeaker system 604. Those skilled in the art will recognize that the assignment of audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k in a virtual venue 600, as illustrated in Figure 7A, is for illustrative purposes only and is not intended to be limiting. Those skilled in the art will recognize that an audio remixing tool can, without departing from the spirit and scope of this disclosure, assign audio voices 704.1-704.m to other virtual loudspeakers 602.1-602.k in the virtual venue 600 to construct a static audio presentation 700.
[0057] The discussion in Figures 7B–7F describes exemplary operations to be performed by the audio remixing tool to assign audio voices 704.1–704.m to virtual loudspeakers 602.1–602.k in a virtual venue 600 as illustrated in Figure 7A. In some embodiments, the exemplary operations described in more detail below in Figures 7B–7D and / or in more detail below in Figures 8A–8B may form or be included in a software remixing toolkit that can be used by the audio remixing tool to construct exemplary audio presentations described herein, such as static audio presentations 700. In these embodiments, the exemplary operations within the software remixing toolkit may be performed independently or in any combination to construct these audio presentations.
[0058] Figure 7B illustrates a simple instrument assignment operation 710 that can be performed by an audio remixing tool to assign audio sounds produced by simple instruments such as percussion instruments, wind instruments, string instruments, and / or electronic instruments to virtual loudspeakers 602.1-602.k in a virtual venue 600. Generally, audio sounds produced by simple instruments, as illustrated in Figure 7B, represent audio sounds that have different characteristics, parameters, and / or attributes from other audio sounds from audio sounds 704.1-704.m. In some embodiments, audio sounds produced by simple instruments have, to give a few examples, pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity that are not similar to other audio sounds from audio sounds 704.1-704.m. In the exemplary embodiment illustrated in Figure 7B, the audio remixing tool can assign audio sounds produced by simple instruments, such as audio sounds 704.1 produced by a microphone, to a virtual loudspeaker 712 from among the virtual loudspeakers 606.1–606.z. In some embodiments, the virtual loudspeaker 712 can be selected from among the virtual loudspeakers 606.1–606.z based on a composite audio program, such as a composite audio program 150 being played by a virtual venue 600. In some embodiments, the virtual loudspeaker 712 can be selected from among the virtual loudspeakers 606.1–606.z through algorithmic best source decoding with respect to the virtual loudspeakers 606.1–606.z. In these embodiments, algorithmic best source decoding will be obvious to those skilled in the art without departing from the spirit and scope of this disclosure. While audio sounds produced by simple instruments are described in Figure 7B as being assigned to virtual loudspeaker 712, those skilled in the art will recognize that audio sounds produced by simple instruments may be assigned to any of the virtual loudspeakers 602.1–602.k in the virtual venue 600 in a manner substantially similar to that described in Figure 7B, without departing from the spirit and scope of this disclosure.In some embodiments, the virtual loudspeaker 712 can be positioned at three-dimensional coordinates (x v , y v , z v ) within the three-dimensional space of the virtual venue 600. As described herein, the three-dimensional coordinates (x v , y v , z v ) can be characterized as providing a reference system or virtual coordinate system for the virtual loudspeaker 712 within the virtual venue 600.
[0059] Figure 7C illustrates in a diagram the assignment operation 720 of a collection of instruments, which can be performed by an audio remixing tool to assign audio sounds produced by a collection of instruments, such as percussion instruments, wind instruments, string instruments, and / or electronic instruments, from instruments of the same classification and / or from instruments of different classifications, to virtual loudspeakers 602.1-602.k in a virtual venue 600. Generally, audio sounds produced by a collection of instruments as illustrated in Figure 7C represent audio sounds that have different characteristics, parameters, and / or attributes from other audio sounds from audio sounds 704.1-704.m. In some embodiments, audio sounds produced by a collection of instruments have, for example, pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity that are not similar to other audio sounds from audio sounds 704.1-704.m. In the exemplary embodiment illustrated in Figure 7C, the audio remixing tool can assign audio sounds produced by a collection of instruments, such as audio sounds 704.2 produced by an acoustic guitar, an electric guitar, and a bass guitar, to a virtual loudspeaker 722 from among the virtual loudspeakers 606.1–606.z. In some embodiments, the virtual loudspeaker 722 can be selected from among the virtual loudspeakers 606.1–606.z based on a composite audio program, such as a composite audio program 150 being played by a virtual venue 600. In some embodiments, the virtual loudspeaker 722 can be selected from among the virtual loudspeakers 606.1–606.z through algorithmic best source decoding with respect to the virtual loudspeakers 606.1–606.z. In these embodiments, algorithmic best source decoding will be obvious to those skilled in the art without departing from the spirit and scope of the present disclosure.The audio sound generated by the ensemble of musical instruments is described as being assigned to the virtual loudspeaker 722 in FIG. 7C. However, one skilled in the art will recognize that the audio sound generated by the ensemble of musical instruments can be assigned to any of the virtual loudspeakers 602.1 - 602.k within the virtual venue 600 in a manner that is substantially similar to that described in FIG. 7C without departing from the spirit and scope of the present disclosure. In some embodiments, the virtual loudspeaker 722 can be positioned at the three-dimensional coordinates (x. G , y G , z G ) within the three-dimensional space of the virtual venue 600. As described herein, the three-dimensional coordinates (x G , y G , z G ) can be characterized as providing a reference system or virtual coordinate system for the virtual loudspeaker 722 within the virtual venue 600.
[0060] Figure 7D illustrates an audio ensemble assignment operation 730 that can be performed by an audio remixing tool to assign audio sounds produced by an audio ensemble having a collection of instruments, such as percussion instruments, wind instruments, string instruments, and / or electronic instruments, and / or instruments from different classifications, to virtual loudspeakers 602.1-602.k in a virtual venue 600. Generally, the audio sounds produced by an audio ensemble represent audio sounds from among audio sounds 704.1-704.m that have similar characteristics, parameters, and / or attributes to each other. In some embodiments, the audio sounds produced by an audio ensemble have, to some examples, similar pitch, volume, timbre, frequency, amplitude, wavelength, and / or velocity to each other. In the exemplary embodiment illustrated in Figure 7D, the audio remixing tool can, in a manner substantially similar to that described herein, assign audio sounds produced by simple instruments, such as audio sound 704.3 produced by a bass drum, from an audio ensemble to a virtual loudspeaker 732 from among the virtual loudspeakers 606.1–606.z. Although the audio sounds produced by simple instruments are described in Figure 7D as being assigned to virtual loudspeaker 712, those skilled in the art will recognize that the audio sounds produced by simple instruments can also be assigned to any of the virtual loudspeakers 602.1–602.k in the virtual venue 600 in a manner substantially similar to that described in Figure 7D, without departing from the spirit and scope of this disclosure. Subsequently, the audio remixing tool can analyze audio sounds 704.1-704.m and identify that another audio sound produced by another simple instrument, such as audio sound 704.m produced by a snare drum, has similar characteristics, parameters, and / or attributes to an audio sound produced by another simple instrument, such as audio sound 704.3 produced by a bass drum.In some embodiments, the audio remixing tool can access an ensemble boundary volume 504, as described herein, which defines the spatial distance between simple instruments in the three-dimensional space of the virtual venue 600. After accessing the ensemble boundary volume 504, the audio remixing tool can, in a manner substantially similar to that described herein, assign audio sounds from the audio ensemble, such as audio sound 704.m produced by a snare drum, to virtual loudspeakers 734, which are spatially positioned within the ensemble boundary volume 504 from among the virtual loudspeakers 606.1-606.z. In some embodiments, virtual loudspeakers 732 and 734 are each located at three-dimensional coordinates (x.) in the three-dimensional space of the virtual venue 600. BD ,y BD ,z BD ) and 3D coordinates (x SD ,y SD ,z SD It can be positioned in 3D coordinates (x BD ,y BD ,z BD ) and 3D coordinates (x SD ,y SD ,z SD ) can be characterized as providing a reference frame or virtual coordinate system with respect to virtual loudspeakers 732 and 734 within the virtual venue 600, respectively. In some embodiments, the audio remixing tool provides the distance between audio sounds produced by simple instruments, such as audio sound 704.3 produced by a bass drum, and audio sounds produced by other simple instruments, such as audio sound 704.m produced by a snare drum, i.e., the distance between the three-dimensional coordinates (x) in the three-dimensional space of the virtual venue 600. BD ,y BD ,z BD ) and 3D coordinates (x SD ,y SD ,z SDThe distance between ) and ) can be estimated. In these embodiments, the audio remixing tool can compare this distance to the ensemble boundary volume 504 to verify that the audio sounds produced by one simple instrument and the audio sounds produced by another simple instrument are within the region of the ensemble boundary volume 504.
[0061] As illustrated herein in Figures 7E and 7F, one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m may influence the assignment of the audio voices 704.1-704.m to the virtual loudspeakers 602.1-602.k within the virtual venue 600. In some embodiments, one or more characteristics, parameters, and / or attributes may indicate the spatial positioning between audio sources that generated the audio sounds 704.1-704.m, such as electronic, mechanical, and / or electromechanical devices, one or more non-human beings, and / or one or more human beings, audio transients of the audio sounds 704.1-704.m, spatial movement of the audio sources that generated the audio sounds 704.1-704.m, timing relationships between the audio sounds 704.1-704.m, audio effects within the audio sounds 704.1-704.m, and / or any other suitable characteristics, parameters, and / or attributes of or between the audio sounds 704.1-704.m, which would be recognized by those skilled in the art without departing from the spirit and scope of this disclosure.
[0062] In some embodiments, the audio remixing tool can access one or more characteristics, parameters, and / or attributes of the audio sounds 704.1-704.m. In these embodiments, the audio remixing tool can analyze the audio sounds 704.1-704.m and identify one or more characteristics, parameters, and / or attributes of these audio sounds in a manner substantially similar to that described herein. Alternatively, or in addition, the audio remixing tool can access a database as described herein and retrieve one or more characteristics, parameters, and / or attributes of the audio sounds 704.1-704.m. After accessing one or more characteristics, parameters, and / or attributes, the audio remixing tool can assign the audio sounds produced by the instrument ensemble to virtual loudspeakers 602.1-602.k in the virtual venue 600 according to these characteristics, parameters, and / or attributes of the audio sounds 704.1-704.m.
[0063] Figure 7E illustrates in a diagram an audio effect operation 740 that can be performed by an audio remixing tool to assign audio sounds produced by simple instruments such as percussion instruments, wind instruments, string instruments, and / or electronic instruments, and / or sets of instruments such as percussion instruments, wind instruments, string instruments, and / or electronic instruments, and / or instruments from different classifications, to give a few examples, to virtual loudspeakers 602.1-602.k in a virtual venue 600. As described herein, the audio remixing tool can access one or more characteristics, parameters, and / or attributes of audio sounds 704.1-704.m, such as audio effects within audio sounds 704.1-704.m, to give a few examples. Alternatively, or in addition, an audio remixing tool may analyze audio voices 704.1–704.m and identify one or more characteristics, parameters, and / or attributes of the audio voices 704.1–704.m in a manner substantially similar to that described herein. In the exemplary embodiment illustrated in Figure 7E, the audio remixing tool may isolate the audio voice, referred to as the parent audio voice, which is produced by a simple instrument and / or a collection of instruments, from the audio effects within the audio voices 704.1–704.m, referred to as the child audio voice. In some embodiments, these audio effects may include, to name a few, one or more echoes, flangers, phasers, choruses, equalizations, filtering, overdrives, pitch shifts, time stretches, resonators, voice effects, synthesizers, modulations, compressions, and / or equivalents. As illustrated in Figure 7E, the audio remixing tool can isolate the audio from the snare drum-generated audio 704.m, referred to as the parent audio 742 in Figure 7E, from the audio effects within the snare drum-generated audio 704.m, referred to as the child audio 744 in Figure 7E.
[0064] After isolating the parent and child audio signals, the audio remixing tool can identify parent-child real-world loudspeaker pairings from among the virtual loudspeakers 602.1-602.k in the virtual venue 600. In some embodiments, the parent-child real-world loudspeaker pairings may include a virtual loudspeaker 746 from among the virtual loudspeakers 606.1-606.z, which is associated with an effect-extended virtual array real-world loudspeaker system 748 from among the effect-extended virtual array real-world loudspeaker systems 608.1-608.l. In some embodiments, the audio remixing tool can identify parent-child real-world loudspeaker pairings by utilizing a predetermined set of source separation rules. In these embodiments, these source separation rules may be based on transients and / or diffuse audio signals within the parent and / or child audio signals. For example, parent-child real-world loudspeaker pairings may be based on diffuse audio signals within the parent and / or child audio signals. In some embodiments, a predetermined set of source separation rules preferably maintains the angular and / or distance relationship between the parent audio and the child audio as much as possible.
[0065] In the exemplary embodiment illustrated in Figure 7E, the audio remixing tool can assign parent audio, generated by a simple instrument and / or collection of instruments such as parent audio 742, to a virtual loudspeaker 746 from among the virtual loudspeakers 606.1-606.z, and the audio remixing tool can assign child audio, generated by a simple instrument and / or collection of instruments such as child audio 744, to an effect-extended virtual array real-world loudspeaker system 748 from among the effect-extended virtual array real-world loudspeaker systems 608.1-608.l. Parent audio sounds produced by simple instruments and / or sets of instruments are described as being assigned to virtual loudspeaker 746, and child audio sounds produced by simple instruments and / or sets of instruments are described as being assigned to the effects-extended virtual array real-world loudspeaker system 748 in Figure 7E. However, those skilled in the art will recognize that parent audio sounds and / or child audio sounds may be assigned to any of the virtual loudspeakers 602.1-602.k in the virtual venue 600 in a manner substantially similar to that described in Figure 7E, without departing from the spirit and scope of this disclosure. In some embodiments, virtual loudspeaker 746 and effects-extended virtual array real-world loudspeaker system 748 are each located in the three-dimensional coordinate system (x) of the three-dimensional space of the virtual venue 600. SDPARENT ,y SDPARENT ,z SDPARENT ) and 3D coordinates (x SDCHILD ,y SDCHILD ,z SDCHILD It can be positioned in 3D coordinates (x SDPARENT ,y SDPARENT ,z SDPARENT ) and 3D coordinates (x SDCHILD ,y SDCHILD ,z SDCHILD These can be characterized as providing a reference frame or virtual coordinate system for the virtual loudspeaker 746 and the effect-extended virtual array real-world loudspeaker system 748, respectively.
[0066] Figure 7F illustrates in a diagram an audio transient operation 750 that can be performed by an audio remixing tool to assign audio sounds produced by a collection of instruments, such as simple instruments like percussion instruments, wind instruments, string instruments, and / or electronic instruments, and / or instruments from the same category of instruments, and / or instruments from different categories, to some examples, to virtual loudspeakers 602.1-602.k in a virtual venue 600. As described herein, the audio remixing tool can access one or more characteristics, parameters, and / or attributes of audio sounds 704.1-704.m, such as audio transients of audio sounds 704.1-704.m, to some example. Alternatively, or in addition, an audio remixing tool may analyze audio voices 704.1–704.m and identify one or more characteristics, parameters, and / or attributes of audio voices 704.1–704.m in a manner substantially similar to that described herein.
[0067] In the exemplary embodiment illustrated in Figure 7E, the audio remixing tool can isolate the attack transient and the decay transient from the audio sounds 704.1–704.m. As illustrated in Figure 7F, the audio remixing tool can isolate the attack transient from the audio sounds 704.m generated by the snare drum, referred to as attack transient audio sound 752 in Figure 7F, and the decay transient from the audio sounds 704.m generated by the snare drum, referred to as decay transient audio sound 754 in Figure 7F. In some embodiments, attack transient audio sound 752 represents a first duration until the snare drum reaches its maximum amplitude, and decay transient audio sound 754 represents a second duration until the snare drum decreases from its maximum amplitude to its steady-state amplitude.
[0068] After isolating the attack and decay transients, the audio remixing tool can identify the attack-decay real-world loudspeaker pairing from among the virtual loudspeakers 602.1-602.k in the virtual venue 600. In some embodiments, the attack-decay real-world loudspeaker pairing may include a virtual loudspeaker 756 from among the virtual loudspeakers 606.1-606.z, which is associated with an effects-extended virtual array real-world loudspeaker system 758 from among the effects-extended virtual array real-world loudspeaker systems 608.1-608.l.
[0069] In the exemplary embodiment illustrated in Figure 7F, the audio remixing tool can assign attack transients generated by simple instruments and / or sets of instruments, such as attack transient audio voice 752, to a virtual loudspeaker 756 from among the virtual loudspeakers 606.1-606.z, and the audio remixing tool can assign decay transients generated by simple instruments and / or sets of instruments, such as decay transient audio voice 754, to an effect-extended virtual array real-world loudspeaker system 758 from among the effect-extended virtual array real-world loudspeaker systems 608.1-608.l. Attack transients generated by simple instruments and / or sets of instruments are described as being assigned to virtual loudspeaker 756, and decay transients generated by simple instruments and / or sets of instruments are described as being assigned to the effect-extended virtual array real-world loudspeaker system 758 in Figure 7F. However, those skilled in the art will recognize that attack audio sounds and / or decay audio sounds may be assigned to any of the virtual loudspeakers 602.1-602.k in the virtual venue 600 in a manner substantially similar to that described in Figure 7F, without departing from the spirit and scope of this disclosure. In some embodiments, virtual loudspeaker 756 and effect-extended virtual array real-world loudspeaker system 758 are located in the three-dimensional coordinate system (x) of the three-dimensional space of the virtual venue 600, respectively. SDATTACK ,y SDATTACK ,z SDATTACK ) and 3D coordinates (x SDDECAY ,y SDDECAY ,z SDDECAY It can be positioned in 3D coordinates (x SDATTACK ,y SDATTACK ,z SDATTACK ) and 3D coordinates (x SDDECAY ,y SDDECAY ,z SDDECAYThese can be characterized as providing a reference frame or virtual coordinate system for the virtual loudspeaker 756 and the effect-extended virtual array real-world loudspeaker system 758, respectively.
[0070] Although not shown in Figures 7A-7F, one or more characteristics, parameters, and / or attributes of a real-world venue, such as the seating arrangement within the real-world venue, the location of the performance stage within the real-world venue, and / or the location of real-world loudspeakers within the real-world venue, may influence the assignment of audio sounds 704.1-704.m to virtual loudspeakers 602.1-602.k within the virtual venue 600. For example, the assignment of audio sounds 704.1-704.m to virtual loudspeakers 602.1-602.k within the virtual venue 600 may be based on time domain, volume level, coverage uniformity, and frequency bandwidth, which are adjusted by an analysis of the artist's original positioning intentions. In some embodiments, the seating arrangement within the real-world venue can determine the degree of influence on the assignment of audio sounds 704.1-704.m to virtual loudspeakers 602.1-602.k within the virtual venue 600. In some embodiments, the assignment of audio sounds 704.1-704.m to virtual loudspeakers 602.1-602.k within a virtual venue 600 can be determined by the creative artist's intent, the transient nature of the content, and the temporal relationship of the content to other content elements. (An exemplary dynamic audio presentation that can be constructed using an exemplary audio demixing tool)
[0071] Figures 8A–8B illustrate the operation of an exemplary audio remixing tool in constructing an exemplary dynamic audio presentation to play a composite audio program on real-world loudspeakers in an exemplary real-world venue, according to some exemplary embodiments of the present disclosure. In the exemplary embodiment illustrated in Figure 8A, an audio remixing tool, such as audio remixing tool 154, can intelligently construct a dynamic audio presentation 800 that can play a composite audio program on real-world loudspeakers in a real-world venue, such as real-world venue 102, as described herein. As described herein, the audio remixing tool can utilize a virtual venue 600, as described herein, and assign the audio sounds of the composite audio program, such as audio sounds 156.1–156.n of the composite audio program 150, as described herein, to real-world loudspeakers in the real-world venue to construct the dynamic audio presentation 800. Those skilled in the art will recognize that the dynamic audio presentation 800 as illustrated in Figure 8A is for illustrative purposes only and is not intended to be limiting. Those skilled in the art will recognize that, without departing from the spirit and scope of this disclosure, audio remixing tools can intelligently construct other dynamic audio presentations and play other composite audio programs on real-world loudspeakers in real-world venues.
[0072] The discussion in Figures 8A-8B describes exemplary operations that may be utilized by an audio remixing tool to construct a dynamic audio presentation 800 for playing a composite audio program on a real-world loudspeaker in a real-world venue. Those skilled in the art will recognize that these operations may be performed independently or in any combination to construct other dynamic audio presentations for playing other composite audio programs on a real-world loudspeaker in a real-world venue, without departing from the spirit and scope of this disclosure.
[0073] In the exemplary embodiment illustrated in Figure 8A, the audio remix can access an electronic library of audio sounds 702 having audio sounds 704.1–704.m present in the composite audio program in a manner substantially similar to that described herein. The audio remixing tool can then assign the audio sounds 704.1–704.m to virtual loudspeakers 602.1–602.k in the virtual venue 600 and construct a dynamic audio presentation 800 for playing the composite audio program on real-world loudspeakers in the real-world venue. As described herein, the virtual venue 600 may include a proscenium virtual loudspeaker system 604 having virtual loudspeakers 606.1–606.z, an effects-extended virtual array real-world loudspeaker system 608.1–608.l, and / or an environment virtual array real-world loudspeaker system 610.1–610.m. As illustrated in Figure 8A, the audio remixing tool can construct a dynamic audio presentation 800 by assigning audio sounds 704.1 produced by a microphone to virtual loudspeaker 606.4 from the proscenium virtual loudspeaker system 604, assigning audio sounds 704.2 produced by an acoustic guitar, electric guitar, and bass guitar to virtual loudspeaker 606.7 from the proscenium virtual loudspeaker system 604, assigning audio sounds 704.3 produced by a bass drum to virtual loudspeaker 606.2 from the proscenium virtual loudspeaker system 604, and / or assigning audio sounds 704.m produced by a snare drum to virtual loudspeaker 606.1 from the proscenium virtual loudspeaker system 604, in a manner substantially similar to that described herein through 7F. Those skilled in the art will recognize that the assignment of audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k in a virtual venue 600, as illustrated in Figure 8A, is for illustrative purposes only and is not intended to be limiting.Those skilled in the art will recognize that an audio remixing tool can, without departing from the spirit and scope of this disclosure, assign audio voices 704.1-704.m to other virtual loudspeakers 602.1-602.k within a virtual venue 600 to construct a dynamic audio presentation 800.
[0074] After assigning audio voices 704.1-704.m to virtual loudspeakers 602.1-602.k, the audio remixing tool can perform a simple spatial movement of the audio voices 704.1-704.m in the three-dimensional space of the virtual venue 600 according to one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m. As described herein, the audio remixing tool can access one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m, for example, the spatial movement of the simple instruments and / or ensembles of instruments that generated the audio voices 704.1-704.m. Alternatively, or in addition, an audio remixing tool may analyze audio sounds 704.1–704.m, for example, audio sounds 704.m produced by a snare drum as illustrated in Figure 8A, and identify one or more characteristics, parameters, and / or attributes of audio sounds 704.1–704.m in a manner substantially similar to that described herein.
[0075] In the exemplary embodiment illustrated in Figure 8A, one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m may indicate that the audio remixing tool performs a simple spatial movement 810 within the three-dimensional space of the virtual venue 600. In some embodiments, the simple spatial movement 810 may include, for example, the spatial movement of the audio voices 704.m produced by a snare drum in the three-dimensional space of the virtual venue 600 from virtual loudspeakers 606.1-606.z to effect-extended virtual array real-world loudspeaker system 608.1-608.l. In some embodiments, one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m may indicate the time t during which the audio voices 704.m produced by the snare drum move from virtual loudspeakers 606.1-606.z to effect-extended virtual array real-world loudspeaker system 608.1-608.l move Those skilled in the art will recognize that the simple spatial movement 810 from the virtual loudspeakers 606.1-606.z to the effect-extended virtual array real-world loudspeaker system 608.1-608.l, as illustrated in Figure 8A, is for illustrative purposes only and is not intended to be limiting. Those skilled in the art will recognize that an audio remixing tool may perform other simple spatial movements to other audio sounds from among the audio sounds 704.1-704.m in a manner substantially similar to the simple spatial movement 810, without departing from the spirit and scope of this disclosure.
[0076] In some embodiments, a simple spatial movement 810 occurs over time t moveIn this context, instantaneous or nearly instantaneous spatial movement, also referred to as snap spatial movement, can be represented in the three-dimensional space of the virtual venue 600 from virtual loudspeakers 606.1-606.z to effect-extended virtual array real-world loudspeaker systems 608.1-608.l. Alternatively, or in addition, simple spatial movement 810 can represent a gradual spatial movement in the three-dimensional space of the virtual venue 600 from virtual loudspeakers 606.1-606.z to effect-extended virtual array real-world loudspeaker systems 608.1, followed by snap spatial movement to effect-extended virtual array real-world loudspeaker systems 608.1-608.l. In some embodiments, snap spatial movement can occur in response to an event, for example, during a gap in audio sound 704.m, such as a period when the snare drum does not produce audio sound. For example, a snap spatial movement could be an instantaneous snap between virtual loudspeakers. In another embodiment, a snap spatial movement could be a slow, gradual snap away from a virtual loudspeaker.
[0077] Figure 8B illustrates in a diagram the complex spatial movement operations 820 that can be performed by an audio remixing tool to assign audio sounds produced by simple instruments such as percussion instruments, wind instruments, string instruments, and / or electronic instruments, and / or sets of instruments such as percussion instruments, wind instruments, string instruments, and / or electronic instruments, and / or instruments from different classifications, to virtual loudspeakers 602.1-602.k in a virtual venue 600. As described herein, the audio remixing tool can isolate attack transients and decay transients from audio sounds 704.1-704.m in a manner substantially similar to that described herein.
[0078] As illustrated in Figure 8B, the audio remixing tool can isolate, in a manner substantially similar to that described herein, the attack transient from the audio sound 704.m produced by the snare drum, referred to as attack transient audio sound 752 in Figure 8A, and the decay transient from the audio sound 704.m produced by the snare drum, referred to as decay transient audio sound 754 in Figure 8A. After isolating the attack transient and decay transient, the audio remixing tool can identify attack-decay real-world loudspeaker pairings from virtual loudspeakers 602.1-602.k in the virtual venue 600, in a manner substantially similar to that described herein. In the exemplary embodiment illustrated in Figure 8B, the audio remixing tool can assign attack transients generated by a simple instrument and / or collection of instruments, such as attack transient audio voice 752, to a virtual loudspeaker 756 from among the virtual loudspeakers 606.1-606.z, and the audio remixing tool can similarly assign decay transients generated by a simple instrument and / or collection of instruments, such as decay transient audio voice 754, to a virtual loudspeaker 756 from among the virtual loudspeakers 606.1-606.z. Attack and decay audio sounds generated by a simple instrument and / or a collection of instruments are described as being assigned to the virtual loudspeaker 756 in Figure 8B, but those skilled in the art will recognize that attack and / or decay audio sounds may be assigned to any of the virtual loudspeakers 602.1–602.k in the virtual venue 600 in a manner substantially similar to that described in Figure 8B, without departing from the spirit and scope of this disclosure.
[0079] After assigning attack and decay audio, the audio remixing tool can perform complex spatial movements of the audio 704.1-704.m in the three-dimensional space of the virtual venue 600 according to one or more characteristics, parameters, and / or attributes of the audio 704.1-704.m. As described herein, the audio remixing tool can access one or more characteristics, parameters, and / or attributes of the audio 704.1-704.m, for example, the spatial movements of the simple instruments and / or ensembles of instruments that generated the audio 704.1-704.m. Alternatively, or in addition, an audio remixing tool may analyze audio sounds 704.1–704.m, for example, audio sounds 704.m produced by a snare drum as illustrated in Figure 8A, and identify one or more characteristics, parameters, and / or attributes of audio sounds 704.1–704.m in a manner substantially similar to that described herein.
[0080] In the exemplary embodiment illustrated in Figure 8B, one or more characteristics, parameters, and / or attributes of the audio voices 704.1-704.m may indicate that the audio remixing tool performs a complex spatial movement 820 within the three-dimensional space of the virtual venue 600. In some embodiments, the complex spatial movement 820 may include, for example, a simple spatial movement 812 of an attack transient audio voice 752 generated by a snare drum in the three-dimensional space of the virtual venue 600 from virtual loudspeakers 606.1-606.z to an effects-extended virtual array real-world loudspeaker system 608.1-608.l, and a simple spatial movement 814 of an attack transient audio voice 752 generated by a snare drum in the three-dimensional space of the virtual venue 600 from virtual loudspeakers 606.1-606.z to an effects-extended virtual array real-world loudspeaker system 608.1-608.l. Those skilled in the art will recognize that the complex spatial movement 820 from virtual loudspeakers 606.1–606.z to the effect-extended virtual array real-world loudspeaker system 608.1, as illustrated in Figure 8B, is for illustrative purposes only and is not intended to be limiting. Those skilled in the art will recognize that an audio remixing tool may perform other simple spatial movements to other audio sounds from among the audio sounds 704.1–704.m in a manner substantially similar to the complex spatial movement 820, without departing from the spirit and scope of this disclosure. In some embodiments, simple spatial movements 812 and / or simple spatial movements 814 may be performed in a manner substantially similar to those described herein. In these embodiments, simple spatial movement 812 may be performed before, simultaneously with, or after simple spatial movement 814. (Exemplary configuration of an exemplary real-world venue for implementing an exemplary audio presentation)
[0081] Figure 9 illustrates, in diagrammatic form, a high-level graphical mapping of an exemplary virtual venue to an exemplary real-world venue in some exemplary embodiments of the present disclosure. In the exemplary embodiments illustrated in Figure 9, an audio remixing tool, such as audio remixing tool 154, can assign audio sounds, such as audio sounds 156.1–156.n, to virtual loudspeakers 602.1–602.k in the virtual venue 600, and construct audio presentations, such as static audio presentations 700 and / or dynamic audio presentations 800, for playing a composite audio program in a real-world venue, such as real-world venue 102. As described herein, the audio remixing tool can configure a real-world venue, as defined in the audio presentation, to play a composite audio program in the real-world venue.
[0082] As illustrated in Figure 9, the audio remixing tool can construct an audio presentation by assigning audio to virtual loudspeakers 602.1-602.k in the three-dimensional space of the virtual venue 600 in a manner substantially similar to that described herein. In some embodiments, the real-world venue may include audio equipment such as amplifiers, crossovers, equalizers, and / or mixers to route and / or signal-tune the audio for playback within the real-world venue. In these embodiments, the audio remixing tool 154 can generate audio control signals such as audio control signals 158.1-158.i and configure the audio equipment in the real-world venue to play the audio presentation on real-world loudspeakers within the real-world venue. In these embodiments, the audio control signals can cause the audio equipment in the real-world venue to route and / or signal-tune the audio for playback through real-world loudspeakers within the real-world venue as defined in the audio presentation.
[0083] As illustrated in Figure 9, the audio remixing tool can access real-world venue configuration information 902, generate audio control signals such as audio control signals 158.1-158.i, and configure audio equipment in the real-world venue to play an audio presentation on real-world loudspeakers within the venue. In some embodiments, the real-world venue configuration information 902 represents an organized collection of data, often referred to as a database. The database may include one or more data tables having data values such as alphanumeric strings, integers, decimals, floating-point numbers, dates, times, binary values, Boolean values, and / or enumerations, to name a few. The database may be a column-oriented database, a relational database, a keystore database, a graph database, and / or a document store, to name a few. In some embodiments, the real-world venue configuration information 902 may include real-world loudspeaker configuration information 904.1-904.k, corresponding to virtual loudspeakers 602.1-602.k installed in the three-dimensional space of the virtual venue 600. In some embodiments, the audio remixing tool can access real-world loudspeaker information 904.1-904.k and generate audio control signals to play audio sounds assigned to virtual loudspeakers 602.1-602.k on real-world loudspeakers in a real-world venue.In these embodiments, the audio remixing tool accesses real-world loudspeaker information 904.1 and generates audio control signals to play the audio assigned to virtual loudspeaker 602.1 on real-world loudspeakers in a real-world venue associated with virtual loudspeaker 602.1, and accesses real-world loudspeaker information 904.2 and generates audio control signals to play the audio assigned to virtual loudspeaker 602.2 on real-world loudspeakers in a real-world venue associated with virtual loudspeaker 602.2 It is possible to generate control signals, access real-world loudspeaker information 904.3, generate audio control signals to play audio assigned to virtual loudspeaker 602.3 on real-world loudspeakers in a real-world venue associated with virtual loudspeaker 602.3, access real-world loudspeaker information 904.k, and generate audio control signals to play audio assigned to virtual loudspeaker 602.k on real-world loudspeakers in a real-world venue associated with virtual loudspeaker 602.k. (Example of the operation of an exemplary audio remixing tool)
[0084] Figure 10 illustrates a flowchart of an exemplary audio remixing tool according to some exemplary embodiments of the present disclosure. The present disclosure is not limited to this description of operation. Rather, it will be apparent to those skilled in the art that other operation control flows are also within the scope and spirit of the present disclosure. The following discussion describes an operation control flow 1000 for intelligently constructing an audio presentation and playing a composite audio program, such as a composite audio program 150, in a real-world venue, such as a real-world venue 102. The operation control flow 1000 can be performed, for example, by an audio remixing tool 154. In some embodiments, the operation control flow 1000 can be performed by one or more computing devices, such as an audio playback server 104.
[0085] In operation 1002, the operation control flow 1000 constructs an audio presentation for playing a composite audio program within a real-world venue. The operation control flow 1000 can utilize a software remixing toolkit to construct the audio presentation in a manner substantially similar to that described herein.
[0086] In operation 1004, the operation control flow 1000 can configure a real-world venue as defined in the audio presentation to play a composite audio program within the real-world venue. The operation control flow 1000 can identify audio control signals, such as audio control signals 158.1-158.i, which configure the real-world venue to play an audio presentation through real-world loudspeakers within the real-world venue in order to play a composite audio program within the real-world venue, in a manner substantially similar to that described herein. (An exemplary computing device that could be used to implement an electronic device in an exemplary real-world venue)
[0087] Figure 11 illustrates a simplified block diagram of a computing device that may be used to implement an electronic device in an exemplary real-world venue according to several embodiments of the present disclosure. The subsequent discussion of Figure 11 describes a computing device 1100 that may be used to implement an audio playback server 104.
[0088] In the embodiment illustrated in Figure 11, the computing device 1100 includes one or more processors 1102. In some embodiments, one or more processors 1102 may include, or be, any of the following: a microprocessor, a graphics processing unit, or a digital signal processor, and their electronic equivalents, such as an application-specific integrated circuit ("ASIC") or a field-programmable gate array ("FPGA"). As used herein, the term “processor” typically refers to a tangible data and information processing device that physically transforms data and information using sequence transformations (also referred to as “operations”). Data and information can be physically represented by electrical, magnetic, optical, or acoustic signals that can be stored, accessed, automatically transferred, combined, compared, or otherwise manipulated by the processor. The term “processor” can refer to a single processor and a multi-core system or multi-processor array, including a graphics processing unit, a digital signal processor, a digital processor, or a combination of these elements. A processor may be an electronic device, for example, comprising a digital logic network (e.g., binary logic) or analog (e.g., an operational amplifier). Processors may also operate within a “cloud computing” environment or as “software as a service” (SaaS) to support the performance of related operations. For example, at least part of an operation may be performed by a collection of processors available in a distributed or remote system, which are accessible via a communication network (e.g., the Internet) and via one or more software interfaces (e.g., application programming interfaces (APIs)).In some embodiments, the computing device 1100 may include an operating system such as Microsoft's Windows®, Sun Microsystems' Solaris®, Apple Computer's MacOS, Linux®, or UNIX®. In some embodiments, the computing device 1100 may also include a basic input / output system (BIOS) and processor firmware. The operating system, BIOS, and firmware are used by one or more processors 1102 to control subsystems and interfaces coupled to one or more processors 1102. In some embodiments, one or more processors 1102 may include Intel's Pentium® and Itanium, Advanced Micro Devices' Opteron and Athlon, and ARM Holdings' ARM processors.
[0089] As illustrated in Figure 11, the computing device 1100 may include a machine-readable medium 1104. In some embodiments, the machine-readable medium 1104 may further include a main random access memory ("RAM") 1106, a read-only memory ("ROM") 1108, and / or a file storage subsystem 1110. The RAM 730 may store instructions and data during program execution, and the ROM 732 may store fixed instructions. The file storage subsystem 1110 provides persistent storage for program and data files and may include a floppy disk drive, a CD-ROM drive, an optical drive, flash memory, or a removable media cartridge, in addition to a hard disk drive, associated removable media.
[0090] The computing device 1100 may further include a user interface input device 1112 and a user interface output device 1114. The user interface input device 1112 may include, to name a few, pointing devices such as alphanumeric keyboards, keypads, mice, trackballs, touchpads, styluses, or graphics tablets; audio input devices such as scanners, touchscreens integrated into displays, speech recognition systems, or microphones; eye-tracking recognition; electroencephalogram pattern recognition; and other types of input devices. The user interface input device 1112 may be connected to the computing device 1100 by wire or wirelessly. Generally, the user interface input device 1112 is intended to include all conceivable types of devices and methods for inputting information into the computing device 1100. The user interface input device 1112 typically allows the user to identify objects, icons, text, and equivalents appearing on several types of user interface output devices, e.g., on a display subsystem. The user interface output device 1120 may include non-visual displays such as display subsystems, printers, fax machines, or audio output devices. The display subsystem may include flat panel devices such as cathode ray tubes (CRTs) and liquid crystal displays (LCDs), projection devices, or other devices for producing visible images, such as virtual reality systems. The display subsystem may also provide non-visual displays via audio output or haptic output (e.g., vibration) devices. In general, the user interface output device 1120 is intended to include all conceivable types of devices and methods for outputting information from the computing device 1100.
[0091] The computing device 1100 may further include a network interface 1116 for providing an interface to an external network, including an interface to a communication network 1118, which is coupled to a corresponding interface device in another computing device or machine via the communication network 1118. The communication network 1118 may comprise many interconnected computing devices, machines, and communication links. These communication links may be wired links, optical links, wireless links, or any other devices for the transmission of information. The communication network 1118 may be any suitable computer network, such as a wide area network like the Internet, and / or a local area network like Ethernet®. The communication network 1118 may be wired and / or wireless, and the communication network may use encryption and decryption methods, such as those available using a virtual private network. The communication network uses one or more communication interfaces that can receive data from and transmit data to other systems. Embodiments of the communication interface typically include Ethernet® cards, modems (e.g., telephone, satellite, cable, or ISDN), (asynchronous) digital subscriber line (DSL) units, Firewire® interfaces, USB interfaces, and equivalents. One or more communication protocols such as HTTP, TCP / IP, RTP / RTSP, IPX, and / or UDP may be used.
[0092] As illustrated in Figure 11, one or more processors 1102, machine-readable media 1104, user interface input devices 1112, user interface output devices 1114, and / or network interfaces 1116 can be coupled together to communicate with each other using a bus subsystem 1120. While the bus subsystem 1120 is schematically shown as a single bus, alternative embodiments of the bus subsystem may use a bus. For example, RAM-based main memory can communicate directly with a file storage system using a direct memory access ("DMA") system. (Conclusion)
[0093] For detailed descriptions, accompanying drawings are used to illustrate exemplary embodiments consistent with this disclosure. The use of “exemplary embodiments” in this disclosure indicates that while the described exemplary embodiments may include certain features, structures, or characteristics, not all exemplary embodiments necessarily include such features, structures, or characteristics. Furthermore, such phrases do not necessarily refer to the same exemplary embodiments. Additionally, any feature, structure, or characteristic described in relation to an exemplary embodiment may be included independently or in any combination with features, structures, or characteristics of other exemplary embodiments, whether expressly described or not.
[0094] The detailed description is not intended to be restrictive. Rather, the scope of this disclosure is defined solely by the following claims and their equivalents. It should be understood that the detailed description section, and not the abstract section, is intended to be used to interpret the claims. The abstract section may describe one or more exemplary embodiments, not the entirety of this disclosure, and is therefore not intended to limit this disclosure and the following claims and their equivalents in any way.
[0095] The exemplary embodiments described herein are provided for illustrative purposes only and are not intended to be limiting. Other exemplary embodiments may be conceivable, and modifications may be made to the exemplary embodiments, while remaining within the spirit and scope of this disclosure. This disclosure is described with the help of feature-building blocks that illustrate the implementation of the specified functions and their relationships. The boundaries of these feature-building blocks are optionally defined herein for the convenience of explanation. Alternative boundaries may be defined, insofar as the specified functions and their relationships are adequately implemented.
[0096] Embodiments of the Disclosure may be implemented in hardware, firmware, software applications, or any combination thereof. Embodiments of the Disclosure may also be implemented as instructions stored on a machine-readable medium that can be read and executed by one or more processors. The machine-readable medium may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing network). For example, the machine-readable medium may include non-transient machine-readable media such as read-only memory (ROM), random-access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, and others. In another embodiment, the machine-readable medium may include transient machine-readable media such as electrical, optical, acoustic, or other forms of propagating signals (e.g., carrier waves, infrared signals, digital signals, etc.). Furthermore, firmware, software applications, routines, and instructions may be described herein as performing certain actions. However, it should be understood that such descriptions are for convenience only, and such actions are actually the result of a computing device, processor, controller, or other device executing the firmware, software application, routine, instruction, etc.
[0097] The detailed description of exemplary embodiments fully reveals the general nature of this disclosure that others may readily modify and / or adapt such exemplary embodiments for use without undue experimentation and without departing from the spirit and scope of this disclosure, by applying the knowledge of those skilled in the art. Such adaptations and modifications are therefore intended to be within the meaning of the exemplary embodiments and their equivalents, based on the teachings and guidance presented herein. It should be understood that the terminology or language used herein is for illustrative purposes only, and not limiting, and therefore should be interpreted by those skilled in the art in light of the teachings herein.
Claims
1. An audio playback server for remixing multiple audio sounds of a composite audio program for playback within a real-world venue, wherein the audio playback server is Memory configured to store audio remixing tools, A processor configured to execute the aforementioned audio remixing tool, wherein the audio remixing tool, when executed by the processor, Accessing a virtual venue corresponding to the aforementioned real-world venue, wherein the virtual venue includes a plurality of virtual loudspeakers corresponding to a plurality of real-world loudspeakers within the venue. Accessing the ensemble boundary volume that defines the maximum spatial distance between the first audio and the second audio of an audio ensemble selected from the aforementioned plurality of audio sounds, Assigning the first audio signal to the first virtual loudspeaker from among the plurality of virtual loudspeakers, Assigning the second audio to a second virtual loudspeaker from among the plurality of virtual loudspeakers that are less than the maximum spatial distance from the first audio, The real-world venue is configured such that multiple audio control signals within the real-world venue are identified, the first audio sound is played in the first real-world loudspeaker from among the multiple real-world loudspeakers corresponding to the first virtual loudspeaker, and the second audio sound is played in the second real-world loudspeaker from among the multiple real-world loudspeakers corresponding to the second virtual loudspeaker. The processor is configured to perform the following: An audio playback server equipped with the following features.
2. The audio playback server according to claim 1, wherein the audio remixing tool, when executed by the processor, further configures the processor to access an electronic library of audio sounds having the plurality of audio sounds decomposed from the composite audio program.
3. When the aforementioned audio remixing tool is executed by the aforementioned processor, The first audio and the second audio are logically grouped to form the audio ensemble, Positioning the first audio sound and the second audio sound within the spatial distance of the ensemble boundary volume, Determining the audible flight time until the first audio signal and the second audio signal arrive at a certain location within the virtual venue. The processor is further configured to perform the following: The audio remixing tool, when executed by the processor, further configures the processor to iteratively position the first audio and the second audio, iteratively determine the audible flight time, and form the ensemble boundary volume. The audio playback server according to claim 1, wherein, with each iteration, the audio remixing tool, when executed by the processor, further configures the processor to adjust the spatial distance in order to determine the maximum spatial distance, such that the difference between the audible times of flight exceeds an audible time of flight threshold.
4. The audio playback server according to claim 1, wherein the audio remixing tool, when executed by the processor, further configures the processor to move the second audio sound from the second virtual loudspeaker to a third virtual loudspeaker that is less than the maximum spatial distance from the first audio sound during the composite audio program.
5. The audio playback server according to claim 4, wherein the movement of the second audio sound from the second virtual loudspeaker to the third virtual loudspeaker comprises a snap spatial movement of the second audio sound.
6. When the aforementioned audio remixing tool is executed by the aforementioned processor, To isolate the attack transient and decay transient of the first audio from the first audio, Assigning the attack transient of the first audio sound to the first virtual loudspeaker, and assigning the decay transient of the first audio sound to a third virtual loudspeaker from among the plurality of virtual loudspeakers that are less than the maximum spatial distance from the first audio sound. The audio playback server according to claim 1, further configured to perform the following:
7. When the aforementioned audio remixing tool is executed by the aforementioned processor, During the composite audio program, the attack transient of the first audio sound is moved from the first virtual loudspeaker to a fourth virtual loudspeaker that is less than the maximum spatial distance from the first audio sound, or During the composite audio program, the decay transient of the first audio sound is moved from the third virtual loudspeaker to the fourth virtual loudspeaker. The audio playback server according to claim 6, further configured to perform the following:
8. A method for remixing multiple audio sounds of a composite audio program for playback in a real-world venue, wherein the method is: The computing device accesses a virtual venue corresponding to the real-world venue, wherein the virtual venue includes a plurality of virtual loudspeakers corresponding to a plurality of real-world loudspeakers within the venue. The computing device accesses an ensemble boundary volume that defines the maximum spatial distance between the first audio and the second audio of an audio ensemble selected from the plurality of audio sounds, The computing device assigns the first audio signal to the first virtual loudspeaker from among the plurality of virtual loudspeakers, The computing device assigns the second audio to a second virtual loudspeaker from among the plurality of virtual loudspeakers that are less than the maximum spatial distance from the first audio, The computing device identifies a plurality of audio control signals within the real-world venue, and configures the real-world venue to play the first audio sound in the first real-world loudspeaker from among the plurality of real-world loudspeakers corresponding to the first virtual loudspeaker, and to play the second audio sound in the second real-world loudspeaker from among the plurality of real-world loudspeakers corresponding to the second virtual loudspeaker. Methods that include...
9. The method according to claim 8, further comprising the computing device accessing an electronic library of audio sounds having the plurality of audio sounds decomposed from the composite audio program.
10. The computing device logically groups the first audio and the second audio to form the audio ensemble. The computing device positions the first audio and the second audio within the spatial distance of the ensemble boundary volume, The computing device determines the audible flight time until the first audio and the second audio arrive at a certain location within the virtual venue. The above positioning is repeated iteratively to determine the audible flight time and to form the ensemble boundary volume. It further includes, The method according to claim 8, wherein, with each iteration, the spatial distance is adjusted so that the difference between the audible flight times exceeds an audible flight time threshold in order to determine the maximum spatial distance.
11. The method according to claim 8, further comprising the computing device moving the second audio sound from the second virtual loudspeaker to a third virtual loudspeaker that is less than the maximum spatial distance from the first audio sound during the composite audio program.
12. The method according to claim 11, wherein the movement of the second audio sound from the second virtual loudspeaker to the third virtual loudspeaker comprises a snap spatial movement of the second audio sound.
13. The computing device isolates the attack transient and decay transient of the first audio from the first audio. The computing device assigns the attack transient of the first audio sound to the first virtual loudspeaker, and assigns the decay transient of the first audio sound to a third virtual loudspeaker from among the plurality of virtual loudspeakers that is less than the maximum spatial distance from the first audio sound. The method according to claim 8, further comprising:
14. The computing device, during the composite audio program, moves the attack transient of the first audio sound from the first virtual loudspeaker to a fourth virtual loudspeaker that is less than the maximum spatial distance from the first audio sound, or The computing device moves the decay transient of the first audio sound from the third virtual loudspeaker to the fourth virtual loudspeaker during the composite audio program. The method according to claim 13, further comprising:
15. A real-world venue for playing a composite audio program, wherein the real-world venue is Multiple real-world loudspeakers in the real-world venue, configured to play multiple audio sounds from the composite audio program, An audio playback server, wherein the audio playback server is This involves accessing a virtual venue corresponding to the aforementioned real-world venue, wherein the virtual venue includes a plurality of virtual loudspeakers corresponding to the plurality of real-world loudspeakers. Accessing the ensemble boundary volume that defines the maximum spatial distance between the first audio and the second audio of an audio ensemble selected from the aforementioned plurality of audio sounds, Assigning the first audio signal to the first virtual loudspeaker from among the plurality of virtual loudspeakers, Assigning the second audio to a second virtual loudspeaker from among the plurality of virtual loudspeakers that are less than the maximum spatial distance from the first audio, The process involves identifying multiple audio control signals within the real-world venue, configuring the real-world venue to reproduce the first audio sound in a first real-world loudspeaker from among the multiple real-world loudspeakers corresponding to the first virtual loudspeaker, and reproducing the second audio sound in a second real-world loudspeaker from among the multiple real-world loudspeakers corresponding to the second virtual loudspeaker. An audio playback server and A real-world venue equipped with these features.
16. The real-world venue according to claim 15, wherein the audio playback server is further configured to access an electronic library of audio sounds having the plurality of audio sounds decomposed from the composite audio program.
17. The aforementioned audio playback server The first audio and the second audio are logically grouped to form the audio ensemble, Positioning the first audio sound and the second audio sound within the spatial distance of the ensemble boundary volume, Determining the audible flight time until the first audio signal and the second audio signal arrive at a certain location within the virtual venue. It is further configured to do the following: The audio playback server is further configured to iteratively position the first audio and the second audio, iteratively determine the audible flight time, and form the ensemble boundary volume. The real-world venue according to claim 15, wherein, with each iteration, the audio playback server is further configured to adjust the spatial distance until the difference between the audible time-of-flight thresholds exceeds the maximum spatial distance.
18. The real-world venue according to claim 15, wherein the audio playback server is further configured to move the second audio sound from the second virtual loudspeaker to a third virtual loudspeaker that is less than the maximum spatial distance from the first audio sound during the composite audio program.
19. The real-world venue according to claim 18, wherein the movement of the second audio sound from the second virtual loudspeaker to the third virtual loudspeaker comprises a snap spatial movement of the second audio sound.
20. The aforementioned audio playback server To isolate the attack transient and decay transient of the first audio from the first audio, Assigning the attack transient of the first audio sound to the first virtual loudspeaker, and assigning the decay transient of the first audio sound to a third virtual loudspeaker from among the plurality of virtual loudspeakers that is less than the maximum spatial distance from the first audio sound, During the composite audio program, the attack transient of the first audio sound is moved from the first virtual loudspeaker to a fourth virtual loudspeaker that is less than the maximum spatial distance from the first audio sound, or During the composite audio program, the decay transient of the first audio sound is moved from the third virtual loudspeaker to the fourth virtual loudspeaker. A real-world venue according to claim 15, further configured to perform the following: