Dynamic audio mixing in multi-wireless speaker environment
By separating and mapping component audio signals in a multi-wireless speaker environment and dynamically adjusting the upmix, it solves the sound quality distortion problem caused by the comb effect in traditional technologies and achieves a more interactive and immersive listening experience.
Patent Information
- Application Number
- CN202380094237.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2025-09-12
AI Technical Summary
In a multi-wireless speaker environment, traditional technologies are prone to undesirable comb effects, resulting in a poor listening experience for users, especially when multiple audio output devices are unevenly distributed in party mode, resulting in distorted volume and sound quality.
A computer-implemented method receives an audio input signal, separates multiple component audio signals, maps them to multiple audio output devices, and dynamically adjusts the upmix to adapt to device location and network changes to generate a suitable sound field.
Effectively reduces or eliminates the comb effect, providing a more interactive and immersive listening experience and ensuring each user receives a personalized audio experience based on location.
Smart Images

Figure CN120642347A_ABST
Abstract
Description
Technical Field
[0001] Various embodiments relate generally to audio systems and, more particularly, to dynamic audio mixing techniques in a multi-wireless speaker environment. Background Art
[0002] With the mobile device ( For example With the proliferation of mobile devices (e.g., smartphones, tablets, etc.), the demand for portable audio output devices (also referred to herein as personal speakers) has also increased. Such audio output devices allow listeners (also referred to herein as users) to have speakers for enjoying audio content on the move that provide better audio quality than the speakers included in the mobile devices. One feature of increasingly popular audio output devices is "party mode," in which multiple audio output devices can be communicatively coupled together to form an ad hoc network of speakers that synchronize together and output audio content as a single speaker system. Party mode can be coordinated and controlled via the mobile device. Content is output from the mobile device to the speaker network.
[0003] Typically, audio output devices support stereo output via one or more left speakers and one or more right speakers (separated by approximately 10-20 cm). While this amount of speaker spacing may be sufficient for a small number of users listening to a single audio output device, problems arise when multiple such audio output devices are used in party mode. In one example, ten or more audio output devices are deployed in party mode and placed at different locations in the listening environment, with each audio output device receiving the same stereo signal. Each audio output device plays the left channel of the stereo signal on its left speaker and the right channel of the stereo signal on its right speaker. At certain locations in the listening environment, a user may hear multiple left-channel audio signals from multiple audio output devices, and may also hear multiple right-channel audio signals from the same and / or different audio output devices. Due to different travel distances and orientations, each left-channel audio signal arrives at the user at different times. Due to this time difference, some portions of the left-channel audio signals may reinforce each other, resulting in an increase in volume, while other portions of the left-channel audio signals may attenuate each other, resulting in a decrease in volume. Likewise, the right channel audio signals arrive at the user at different times, potentially at different times than the left channel audio signals. Consequently, the user may perceive some portions of the left channel audio as being lower or higher in volume than others in the right channel audio. This phenomenon, known as the combing effect, can lead to an undesirable listening experience.
[0004] As previously stated, there is a need for more efficient techniques for generating audio signals for output by an audio system having multiple audio output devices. Summary of the Invention
[0005] Various embodiments of the present disclosure describe a computer-implemented method for generating an audio signal in an audio system. The method includes receiving an audio input signal. The method also includes separating a plurality of component audio signals from the audio input signal. The method also includes, for a subset of component audio signals included in the plurality of component audio signals, mapping each component audio signal in the subset of component audio signals to one or more of a plurality of audio output devices. The method also includes transmitting each component audio signal to the corresponding one or more audio output devices based on the mapping. The audio signal in the audio system is included on a first computing device.
[0006] Other embodiments provide, among other things, one or more non-transitory computer-readable media and systems configured to implement the methods set forth above.
[0007] At least one technical advantage of the disclosed technology over the prior art is that it enables the deployment of multiple audio output devices (such as personal speakers) in party mode without generating the undesirable combing effect common with conventional techniques. Another technical advantage of the disclosed technology over the prior art is that each user within an environment can have a different audio experience based on their position relative to the multiple audio output devices. Furthermore, the upmixes transmitted to the multiple audio output devices dynamically change based on the speakers in the network, the source audio being used, and so on. More specifically, the technology dynamically changes the upmixes as audio output devices move within the environment, leave the network, or new audio output devices enter the network. In this way, the upmixes are adapted to generate an appropriate sound field as the audio output devices and source audio change over time. Consequently, users can enjoy a more interactive and immersive listening experience compared to conventional techniques. These technical advantages provide one or more technical improvements over prior art approaches. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order that the manner in which the above-described features of various embodiments can be understood in detail, a more particular description of the inventive concepts briefly summarized above may be rendered by reference to various embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the drawings illustrate only typical embodiments of the inventive concepts and are therefore not to be considered limiting of scope, and that other equally effective embodiments may exist.
[0009] Figure 1 is a block diagram of a computer system configured to implement one or more aspects of various embodiments;
[0010] Figure 2 A coordinated audio system according to one or more aspects of various embodiments is shown;
[0011] Figure 3 An example of a listening environment for coordinating an audio system according to one or more aspects of various embodiments is shown;
[0012] Figure 4 Another example of a listening environment for a coordinated audio system according to one or more aspects of various embodiments is shown;
[0013] Figure 5 Yet another example of a listening environment for a coordinated audio system according to one or more aspects of various embodiments is shown; and
[0014] Figure 6 is a flow chart of method steps for generating a set of audio streams for a coordinated audio system according to one or more aspects of various embodiments. DETAILED DESCRIPTION
[0015] In the following description, numerous specific details are set forth to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that the inventive concept may be practiced without one or more of these specific details.
[0016] Figure 1 1 is a diagram illustrating a computer system 100 configured to implement one or more aspects of various embodiments. As shown, the computer system 100 includes, but is not limited to, a computing device 180, an input device 122, an output device 124, an audio output device 126, a network 160, an audio device network 162, and a media content service 170. The computing device 180 includes, but is not limited to, one or more processing units 102, an I / O device interface 104, a network interface 106, an interconnect 112 ( For example , bus), storage device 114, and memory 116. Memory 116 stores, but is not limited to, an output device manager application 150 and an audio upmixing application 152. Processing unit 102 and memory 116 may be implemented in any technically feasible manner. For example, but not limited to, in various embodiments, any combination of processing unit 102 and memory 116 may be implemented as a standalone chip or as part of a more comprehensive solution implemented as an application specific integrated circuit (ASIC), a system on a chip (SoC), etc. Processing unit 102, I / O device interface 104, network interface 106, storage device 114, and memory 116 may be communicatively coupled to one another via interconnect 112.
[0017] The one or more processing units 102 may include any suitable processor, such as a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a tensor processing unit (TPU), any other type of processing unit, or a combination of multiple processing units, such as a CPU configured to operate in conjunction with a GPU. In general, each of the one or more processing units 102 may be any technically feasible hardware unit capable of processing data and / or executing software applications and modules.
[0018] The storage device 114 may include non-volatile storage for applications, software modules, and data, and may include fixed or removable disk drives, flash memory devices, and CD-ROM, DVD-ROM, Blu-ray, HD-DVD, or other magnetic, optical, solid-state storage devices, etc. The storage device 114 may be located in whole or in part in a remote storage system, referred to herein as the “cloud,” and accessed through a connection such as the network 160.
[0019] The memory 116 may include random access memory (RAM) modules, flash memory cells, or any other type of memory cells, or a combination thereof. One or more processing units 102, I / O device interface 104, and network interface 106 are configured to read data from and write data to the memory 116. The memory 116 includes various software programs and modules ( For example , operating system, one or more application programs) and application data associated with the software program ( For example , data loaded from storage device 114).
[0020] In some embodiments, one or more databases 142 are loaded from storage device 114 into memory 116. Databases 142 include applications associated with one or more applications executable by processing unit 102, user data, media content, etc.
[0021] In some embodiments, computing device 180 is communicatively coupled to one or more networks 160. Network 160 may be a network that allows computing device 180 to communicate with other systems or devices, such as servers, cloud computing systems, or other networked computing devices or systems. For example , media content service 170)). For example, network 160 may include a wide area network (WAN), a local area network (LAN), a wireless network ( For example, Wi-Fi network, cellular data network, ad hoc network) and / or the Internet, etc. The computing device 180 can be connected to the network 160 via the network interface 106. In some embodiments, the network interface 106 is hardware, software, or a combination of hardware and software that is configured to connect to the network 160 and interface with the network. In some embodiments, the network interface 106 facilitates communication via one or more standard and / or proprietary protocols ( example like , Bluetooth, proprietary protocols associated with specific manufacturers, etc.) to communicate with other devices or systems.
[0022] The media content service 170 includes a device configured to For example , provides ( For example , one or more computerized services that distribute media content. Examples of media content services 170 include Spotify, Apple Music, Pandora, YouTube Music, Tidal, etc. As used herein, media or media content includes but is not limited to audio content ( For example , spoken and / or musical audio content, audio content files, streaming audio content, audio tracks of videos, etc.) and / or video content. Examples of media content services 170 include, but are not limited to, media content streaming services, YouTube, digital media content vendors, media servers (local and / or remote), etc. More generally, media content services 170 include one or more computer systems ( For example , servers, cloud computing systems, networked computing systems, distributed computing systems, etc.). Computing device 180 may be communicatively coupled with media content service 170 via network 160 and download and / or stream media content from media content service 170.
[0023] Input device 122 includes one or more devices capable of providing input. Examples of input device 122 include, but are not limited to, a touch-sensitive surface ( For example , touchpad), microphone, touch-sensitive screen, buttons, knobs, dials, keyboard, pointing device ( example like , mouse), etc. Output device 124 includes one or more devices that provide output. Examples of output device 124 include, but are not limited to, display devices, tactile devices, etc. Examples of display devices include, but are not limited to, LCD displays, LED displays, OLED displays, AMOLED displays, touch-sensitive displays, transparent displays, projection systems, etc. Additionally, input device 122 and / or output device 124 may include devices that can both receive input and provide output, such as touch-sensitive displays.
[0024] Audio output device 126 ( For example , audio output devices 126-1, 126-2, ... 126-N) include one or more devices capable of outputting sound. Audio output devices 126 include, but are not limited to, portable speakers, bone conduction speakers, shoulder-worn and shoulder-mounted headphones, neck speakers, etc. In some embodiments, audio output devices 126 can be transmitted via I / O device interface 104 and / or network interface 106 in any technically feasible manner ( For example , Universal Serial Bus (USB), Bluetooth, ad hoc Wi-Fi) are coupled to the computing device 180 wired or wirelessly.
[0025] In various embodiments, the audio output device 126 also includes computing, communication, and / or networking capabilities. For example, the audio output device 126 may also include one or more processing units similar to the processing unit 102, memory and / or storage devices, and a network interface similar to the network interface 106. The audio output device 126 may be communicatively coupled to one or more other audio output devices 126 and / or to the computing device 180 via the network interface and optionally store data.
[0026] In various embodiments, multiple audio output devices 126 can be communicatively coupled to each other and / or to the computing device 180 to form a coordinated audio system. As used herein, a coordinated audio system is an ad hoc network of audio output devices 126 that are communicatively coupled to each other and to the computing device 180 via an audio device network 162. The audio device network 162 is typically a wireless network, such as a Wi-Fi network, an ad hoc Wi-Fi network, a Bluetooth network, etc. In some embodiments, the audio output devices 126 in the coordinated audio system operate in "party mode."
[0027] In a coordinated audio system, audio output device 126 outputs media content received from computing device 180 via audio device network 162. Audio output device 126 outputs media content in a synchronized or nearly synchronized manner. In some embodiments, a coordinated audio system is initiated from computing device 180 via output device manager application 150. For example, when computing device 180 is communicatively coupled to audio output device 126, computing device 180 may send media content items to audio output device 126 for synchronized output by audio output device 126. Figure 2 Described in more detail in .
[0028] Memory 116 includes an output device manager application 150 and one or more audio upmixing applications 152. The output device manager application 150 and the audio upmixing applications 152 are stored in memory 116 and loaded into the memory from storage device 114. In some examples, the audio upmixing applications 152 may also be loaded from the cloud and / or executed in the cloud rather than being executed locally on processing unit 102. In operation, the audio upmixing applications 152 output ( For example , decoded for playback) locally stored media content ( example like , stored in memory 114) and / or media content from media content service 170. The audio upmixing application 152 may also communicate with media content service 170 to obtain ( For example , purchase, rent and / or subscribe to download, stream or otherwise retrieve) media content for output and / or storage at computing device 180.
[0029] Output device manager application 150 facilitates management of audio output devices 126. In some examples, part or all of output device manager application 150 may be executed on one or more audio output devices 126. When computing device 180 is communicatively coupled to audio output device 126, a user may perform management functions of audio output device 126 via output device manager application 150, including but not limited to monitoring the status of audio output device 126 (e.g., For example , battery level, volume level, firmware version, etc.), configure the settings of the audio output device 126, update the firmware of the audio output device 126, etc. In some embodiments, the output device manager application 150 can ( For example , interfacing with the audio upmixing application 152 via an application programming interface (API) to cause or otherwise facilitate the audio upmixing application 152 to send media content to the audio output device 126 and / or obtain data associated with the audio upmixing application 152, including but not limited to media content library information.
[0030] Furthermore, in some embodiments, the output device manager application 150 facilitates the creation and management of coordinated audio systems. A user can, via the output device manager application 150, command an audio output device 126 to join a coordinated audio system, thereby creating a coordinated audio system. In some examples, an audio output device 126 that was previously part of a coordinated audio system can automatically rejoin the coordinated audio system by being powered on. In some examples, a nearby audio output device 126 can automatically join the coordinated audio system when the nearby audio output device 126 is in communication proximity or powered on near the coordinated audio system. Within the output device manager application 150, a user can configure individual audio output devices 126 and / or configure the coordinated audio system, modify ( For example , add or remove audio output device 126) or terminate the coordinated audio system, and perform other management functions associated with the coordinated audio system. In addition, the output device manager application 150 can generate a media content playlist to be sent by the audio upmixing application 152 (whose output is sent to the audio output device 126).
[0031] In operation, the audio upmixing application 152 receives an audio input signal. In some examples, the audio upmixing application 152 receives the audio input signal from the input device 122. In some examples, the audio upmixing application 152 retrieves data representing the audio input signal from the database 142 stored in the memory 116, from the storage device 114, or the like. Typically, the input audio signal is a stereo audio signal comprising a left audio channel and a right audio channel. Alternatively, the audio input signal is a mono audio signal having a single channel. In some examples, the audio input signal is a multi-channel encoded audio signal comprising more than two channels, such as four channels, six channels, eight channels, or the like. These multi-channel encoded audio signals include four-channel audio, DVD-Audio, Super Audio CD, Dolby Atmos, and the like. Regardless of the format of the input audio signal, the audio upmixing application 152 generates a customized and dynamic upmix comprising multiple component audio signals and transmits these component audio signals to the audio output device 126, as described herein. In some examples, the audio upmixing application 152 executes on cloud computing resources, and the output of the audio upmixing application 152 is transmitted to the computing device 180 and then to the audio output device 126. Additionally or alternatively, the output of the audio upmixing application 152 bypasses the computing device 180 and is transmitted directly to the audio output device 126. In this manner, the audio upmixing application 152 generates a more immersive acoustic environment, referred to as a sound field, when streaming and transmitting audio signals to multiple audio output devices 126. Furthermore, the audio upmixing application 152 generates audio output signals that reduce or eliminate undesirable combining artifacts associated with conventional techniques. In this regard, the audio upmixing application 152 can generate the component audio signals in any technically feasible manner.
[0032] One way to mitigate this comb effect is for the audio upmixing application 152 to transmit the left-channel audio signal to an audio output device positioned on one side of the listening environment, such that the left-channel audio signal is sent to both the left and right speakers of the audio output device. Similarly, the audio upmixing application 152 transmits the right-channel audio signal to another audio output device positioned on the other side of the listening environment, such that the right-channel audio signal is sent to both the left and right speakers of the audio output device. While this technique can reduce the comb effect, it is not suitable for systems that include more than two audio output devices.
[0033] Additionally or alternatively, to generate a customized upmix for the audio output device 126, the audio upmixing application 152 performs a signal separation process (also referred to herein as a blind source separation process) on the audio input signal. The audio upmixing application 152 may perform other source separation techniques known to those skilled in the art to derive component audio signals from the audio input signal. Signal separation is functionally equivalent to or otherwise referred to as source separation, blind signal separation, or blind source separation. Various methods may include principal component analysis, independent component analysis, independent vector analysis, non-negative matrix factorization, independent low-rank matrix analysis, and the like. Any of the various source separation techniques may be performed using classical signal processing techniques or using machine learning techniques, deep learning techniques, artificial intelligence methods, and the like. In some examples, signal separation separates a desired audio signal from an audio input signal that also includes additional signals. In this regard, signal separation can be employed in mobile phones to reduce or eliminate background noise, such as HVAC noise, traffic sounds, background speech, and the like, while preserving the voice audio of the mobile phone user. Additionally, signal separation can be employed in hearing aids to preserve the voice audio from one person and reduce or eliminate the voice audio of others nearby. The audio upmixing application 152 performs a signal separation process on an audio input signal (such as a stereo audio signal) to separate a plurality of component audio signals from the audio input signal.
[0034] In some examples, the audio input signal includes a song performed by a rock band consisting of a vocalist, a lead guitar, a bass guitar, and drums. The audio upmixing application 152 performs a signal separation process on the audio input signal to separate the audio input signal into four component audio signals, one for each of the vocalist, the lead guitar, the bass guitar, and the drums. In some examples, the audio input signal includes a song performed by a ten-piece jazz band consisting of various horns and drums as well as the vocalist. The audio upmixing application 152 performs a signal separation process on the audio input signal to separate the audio input signal into eleven component audio signals, one for the vocalist and one for each instrument in the ten-piece band. In some examples, the audio input signal includes a song performed by a pop band consisting of a lead vocalist, two backing vocalists, a lead guitar, a keyboard, and drums. The audio upmixing application 152 performs a signal separation process on the audio input signal to separate the audio input signal into seven component audio signals, one for each of the lead vocalist, two backing vocalists, the lead guitar, the bass guitar, the keyboard, and drums.
[0035] Some component audio signals are more difficult for the source separator to separate completely and accurately, such as when the component audio signals have very similar characteristics. In some examples, the audio upmixing application 152 can merge two backup singers or all three singers into one audio signal. In some examples, the audio upmixing application 152 can merge two guitars into one audio signal. In some examples, for a drum kit, the audio upmixing application 152 can merge the individual toms together into one component audio signal, or the audio upmixing application 152 can separate each tom into a separate signal. In addition, the audio upmixing application 152 can merge cymbals together into one component audio signal, or the audio upmixing application 152 can combine cymbals with other individual drums in various ways to form any number of component audio signals for a drum kit.
[0036] For each component audio signal, the audio upmixing application 152 maps the component audio signal to one or more audio output devices 126 in any technically feasible combination. In some examples, the audio upmixing application 152 maps each component audio signal to a different audio output device in a one-to-one mapping. As a result, each audio output device plays a different component audio signal. In some examples, the audio upmixing application 152 maps two or more component audio signals to the same audio output device in a two-to-one or many-to-one mapping. As a result, the audio output device 126 plays a mix of the two or more component audio signals. In some examples, the audio upmixing application 152 maps the component audio signals to two or more audio output devices in a one-to-two or one-to-many mapping. As a result, multiple audio output devices can play all or part of the same component audio signal. The audio upmixing application 152 may employ these mapping techniques in any technically feasible combination within the scope of the present disclosure.
[0037] The audio upmixing application 152, in conjunction with the output device manager application 150, transmits each component audio signal to a corresponding audio output device 126 based on the mappings described herein. In some examples, the audio upmixing application 152 assigns the component audio signals to the audio output devices 126 in a random manner.
[0038] In some examples, the audio upmixing application 152 matches the frequency bandwidth and relative volume level of each component audio signal to the frequency bandwidth and maximum loudness of each audio output device 126. In this regard, the audio upmixing application 152 may match a component audio signal including low-frequency audio (such as a bass drum, bass guitar, baritone saxophone, tuba, baritone voice, etc.) to an audio output device 126 that is adapted to reproduce low-frequency audio. Similarly, the audio upmixing application 152 may match a component audio signal including mid-frequency audio (such as a lead guitar, French horn, tenor saxophone, snare drum, alto voice, etc.) to an audio output device 126 that is adapted to reproduce mid-frequency audio. Likewise, the audio upmixing application 152 may match a component audio signal including high-frequency audio (such as a trumpet, hi-hat, soprano voice, etc.) to an audio output device 126 that is adapted to reproduce high-frequency audio.
[0039] In some examples, the audio upmixing application 152 assigns component audio signals to the audio output devices 126 based on the size of the audio output devices 126 and the volume levels of the component audio signals. The audio upmixing application 152 maps component audio signals that include loud (i.e., high) volume audio to audio output devices 126 that are well-suited for reproducing loud volume audio. Similarly, the audio upmixing application 152 maps component audio signals that include soft volume audio to audio output devices 126 that are well-suited for reproducing soft volume audio, such as those that are not capable of playing at high audio output levels. The volume levels that a particular audio output device 126 is well-suited to accommodate can be predetermined from the product specifications of the audio output devices 126, from a user interface that receives user input regarding various audio output devices 126, from metadata associated with the audio output devices 126, from measurements of frequency responses after generating one or more audio frequency sweeps, and the like.
[0040] In some examples, the audio upmixing application 152 assigns the component audio signals to the audio output devices 126 based on the channel assignments received from the graphical user interface. The graphical user interface models the relative positions of the audio output devices 126 in the environment. In some examples, the graphical user interface can be accessed on the computing device 180, the input device 122, via the I / O device interface 104, or the like. In some examples, the graphical user interface provides a phantom image, whereby the audio upmixing application 152 generates and transmits the component audio signals to two or more audio output devices to create the appearance that a particular component audio signal is located at a particular location without the audio output device, even though the original image positions of the component audio signals were located at different locations when the audio input signals were captured or when the audio input signals were mixed and mastered. Some types of recordings, such as classical music recordings and choral recordings, are recorded on a soundstage in the presence of a full orchestra and / or choir. Other types of recordings, such as pop music and rock music, are composed of a combination of various studio recordings of individual musicians and singers that were previously recorded in separate acoustic environments. The various recordings are then mixed and mastered to simulate the sound of the entire band, as if the band were performing on the soundstage. Thus, each of the various studio recordings is associated with a different component audio signal at a different virtual location, where the virtual location is determined when the various studio recordings are mixed and mastered into the final recording. In some examples, the audio upmixing application 152 assigns the component audio signals to the audio output devices 126 based on a mapping between the location of the sounds captured from the soundstage at the time of the recordings relative to the location of the audio output devices 126 in the current environment. In some examples, the audio upmixing application 152 assigns the audio components present in the audio source to the audio output devices 126 based at least in part on metadata included in the audio source. The metadata may include information identifying the individual audio components present in the audio source. The audio upmixing application 152 assigns the individual audio components to the audio output devices 126 based on the metadata.
[0041] In some examples, the audio upmixing application 152 tracks the audio output device 126 and adjusts the component audio signals or mappings of the component audio signals as the audio output device 126 moves within the environment, leaves the environment, enters the environment, etc.
[0042] In some examples, the various components of the coordinated audio system perform the upmix. In some examples, the upmix is generated locally on the computing device 180. Additionally or alternatively, the audio upmix application 152 transmits the audio input signal to one or more audio output devices 126. In such an example, each audio output device 126 can perform a localized audio upmix based on at least one of the stereo audio channels received by the audio output device 126, the mono audio channels received by the audio output device 126, or all audio channels sent to each of the audio output devices. In some examples, the various components of the coordinated audio system perform the upmix via a cloud connection from each audio output device 126, from a designated audio output device 126, from the computing device 180, and the like.
[0043] In some examples, the audio upmixing application 152 transmits the component audio signals to each audio output device 126. Additionally or alternatively, the audio upmixing application 152 transmits the component audio signals to a specific audio output device 126. The specific audio output device 126, in turn, transmits the component audio signals to each of the other audio output devices 126. Additionally or alternatively, the audio upmixing application 152 acts as a remote control unit, where the audio signals are transmitted directly to each audio output device 126, or to all audio output devices 126.
[0044] In some examples, the audio upmixing application 152 generates component audio signals, where each component audio signal represents a different instrument or voice. Additionally or alternatively, the audio upmixing application 152 generates component audio signals, where each component audio signal represents a different region of the sound field, either where the musicians were located during the recording or where the musicians were evidently located during the mixing and mastering process of the recording. Additionally or alternatively, the audio upmixing application 152 generates component audio signals, where each component audio signal represents a different region of the sound field, such as an angular region of the sound field.
[0045] In some examples, the audio upmixing application 152 omits component audio signals in "karaoke" mode so that they are not transmitted to the audio output device 126. In this regard, the user can select one or more component audio signals (such as instruments, voices, or groups of instruments or voices) to be omitted via an input device (such as an interactive graphical user interface implemented on a touch screen). For example, the audio upmixing application 152 can omit one or more component audio signals comprising a vocal part, allowing the user to sing along with the audio. In some examples, the audio upmixing application 152 can omit one or more component audio signals comprising a lead guitar or bass guitar, allowing the user to play guitar along with the audio. In some examples, the audio upmixing application 152 can omit one or more component audio signals comprising drums, allowing the user to play drums along with the audio. In some examples, the audio upmixing application 152 can omit one or more component audio signals representing various audio processing noises, background noises, artificial or natural room reverberations, or artifacts of the source separation process. Various other combinations are possible within the scope of the present disclosure.
[0046] In some examples, the audio upmixing application 152 can guide the user, via an interactive graphical user interface, in positioning the audio output device 126 to achieve accurate sound field reproduction. In this regard, the audio upmixing application 152 can position the component audio signal of the tuba player near the component audio signal of the brass instrument. In this way, the audio upmixing application 152 can match the positions of these instruments during an orchestral or brass band recording. In some examples, the arrangement of the component audio signals can differ in whole or in part from their positions during the orchestral recording, or from their original positions on the sound field resulting from the original recording, mixing, or mastering.
[0047] In some examples, the audio upmixing application 152 generates a room reverberation effect that is inserted into each of the component audio signals to increase the sense of immersion experienced by the user when listening to the audio. To produce this room reverberation effect, once the audio output devices 126 are in place, the audio upmixing application 152 can optionally be informed by room measurements determined by each audio output device 126 based on the 3D position and / or 3D orientation of the audio output device 126. Thus, so-called "dry" rooms can benefit from more additional reverberation, while "live" rooms can benefit from less additional reverberation. In some examples, the audio upmixing application 152 generates the room reverberation effect based on input audio received from microphones placed on or near one or more audio output devices 126.
[0048] Figure 21 shows a coordinated audio system 200 according to one or more aspects of various embodiments. The coordinated audio system 200 includes a plurality of audio output devices 126-1 through 126-N (126-1 through 126-N) communicatively coupled together via an audio device network 162. For example , speakers). The coordinated audio system 200 also includes computing devices 180 communicatively coupled together via the audio device network 162. Within the coordinated audio system 200, the computing devices 180 are communicatively coupled to zero or more audio output devices 126.
[0049] Communication in the coordinated audio system 200 may use standard and / or proprietary protocols. For example, the computing device 180 may use a standard protocol ( For example , Bluetooth, Wi-Fi) to communicate with the audio output device 126 and each other, and the audio output device 126 may use a standard or proprietary protocol ( For example , Bluetooth, a proprietary protocol associated with a specific manufacturer) communicate with each other. In some embodiments, the audio output devices 126 that communicate with each other using a proprietary protocol are audio output devices 126 from the same manufacturer ( For example , speakers of the same brand).
[0050] In the coordinated audio system 200, the computing device 180 receives an audio input signal and separates a plurality of component audio signals from the audio input signal. For each component audio signal included in the plurality of component audio signals, the computing device 180 maps the component audio signal to one or more audio output devices 126. The computing device 180 transmits each component audio signal to a corresponding audio output device 126 based on the mapping. The audio output device 126 outputs audio corresponding to the component audio signal from the audio input signal received from the computing device 180. In some embodiments, data corresponding to the component audio signal can be transmitted from the computing device 180 to at least one audio output device 126, and the data can be transmitted between the audio output devices 126.
[0051] Figure 3An example of a listening environment 300 for coordinating audio systems according to one or more aspects of various embodiments is shown. As shown, the listening environment 300 includes four audio output devices 310, 312, 320, and 330. The four audio output devices 310, 312, 320, and 330 play component audio signals corresponding to drums 340, bass guitar 352, lead guitar 350, and vocals / microphone 370, respectively. Thus, the audio output devices 310, 312, 320, and 330 are mapped to the component audio signals in a one-to-one relationship. The listening environment 300 generated by the audio output devices 310, 312, 320, and 330 provides different audio experiences depending on the user's position within the listening environment 300. As shown, the user 380 is centrally located between the audio output devices 310, 312, 320, and 330. As a result, the user 380 hears a balanced mix of the component audio signals for the drums 340, bass guitar 352, lead guitar 350, and vocals / microphone 370. User 382 is located near audio output device 312. Therefore, user 382 hears an increase in the level of the component audio signal of bass guitar 352. User 384 is located near audio output device 330. Therefore, user 384 hears an increase in the level of the component audio signal of vocals 370. User 386 is located near audio output device 310. Therefore, user 386 hears an increase in the level of the component audio signal of drums 340. The closer a listener is to a particular audio output device, the more dominant that particular component audio signal becomes in the overall sound heard by the listener. The farther a listener is from a particular audio output device, the less the output of that particular device contributes to the overall sound heard by the listener.
[0052] In some examples, the audio upmixing application 152 assigns component audio signals to audio output devices based on matching the frequency spectrum of the component audio signals with the frequency operating range of each audio output device. In this regard, some audio output devices have frequency responses that prevent them from playing the lowest notes at the same output level as mid- and high-frequency notes. Therefore, these audio output devices are not well suited for reproducing signals from instruments with high-amplitude, low-frequency notes. As shown in the figure, drums 340 and bass guitar 352 (which generate audio with relatively high amplitude in the low-frequency spectrum) are assigned to audio output devices 310 and 312, respectively. Audio output devices 310 and 312 are capable of operating at high amplitude within this low-frequency spectrum. Lead guitar 350, which primarily generates output in the mid-frequency spectrum and less output in the low-frequency range, is assigned to audio output device 320. Audio output device 320 has a frequency operating range within this mid-frequency spectrum and is unable to output high-amplitude sounds in the low-frequency range. Vocals / microphone 370, which generates audio with a relatively high frequency spectrum, is assigned to audio output device 330. The audio output device 330 has a frequency operating range within the high frequency spectrum and cannot play loudly in the mid-frequency and low-frequency ranges.
[0053] Figure 4 Another example of a listening environment 400 for coordinating audio systems according to one or more aspects of various embodiments is shown. As shown, the listening environment 400 includes eleven audio output devices 410, 412, 414, 420, 422, 424, 426, 430, 432, 434, and 436. The eleven audio output devices 410, 412, 414, 420, 422, 424, 426, 430, 432, 434, and 436 play component audio signals corresponding to various instruments and voices. Audio output device 410 primarily plays a component audio signal corresponding to baritone saxophone 460. Audio output device 412 primarily plays a component audio signal corresponding to bass drum 444. Audio output device 414 primarily plays a component audio signal corresponding to tuba 456.
[0054] Audio output device 420 primarily plays the component audio signal corresponding to snare drum 440. Audio output device 422 primarily plays the component audio signal corresponding to tenor saxophone 462. Audio output device 424 primarily plays the component audio signal corresponding to snare drum 442. Audio output device 426 primarily plays the component audio signal corresponding to French horn 454. Audio output device 430 primarily plays the component audio signal corresponding to trumpet 450. Audio output device 432 primarily plays the component audio signal corresponding to trombone 452. Audio output device 434 primarily plays the component audio signal corresponding to vocal part / microphone 470. Audio output device 436 primarily plays the component audio signal corresponding to hi-hat 446.
[0055] The listening environment 400 generated by audio output devices 410, 412, 414, 420, 422, 424, 426, 430, 432, 434, and 436 provides different audio experiences depending on the user's location within the listening environment 400. As shown, user 480 is centrally located between eleven audio output devices 410, 412, 414, 420, 422, 424, 426, 430, 432, 434, and 436. Therefore, user 480 hears a balanced mix of the audio signals for all instruments and the vocal part. User 482 is located near audio output devices 410, 422, and 424. Therefore, user 482 primarily hears the component audio signals for baritone saxophone 460, tenor saxophone 462, and snare drum 442, along with lower levels of each of the remaining component audio signals. User 484 is located near audio output devices 426 and 434. Thus, user 484 primarily hears the component audio signals for French horn 454 and lead vocals 470, and each of the remaining component audio signals at a lower level.
[0056] In some examples, the audio upmixing application 152 assigns the component audio signals to the audio output devices based on matching the frequency spectrum of the component audio signals with the frequency operating range of each audio output device. In this regard, some audio output devices have a frequency response that makes them unable to play the lowest notes at the same output level as the mid-frequency and high-frequency notes. Therefore, these audio output devices are not well suited for reproducing signals from instruments with high-amplitude low-frequency notes. As shown in the figure, baritone saxophone 460, bass drum 444, and tuba 456 (which generate audio with relatively high amplitude in the low-frequency spectrum) are assigned to audio output devices 410, 412, and 414, respectively. Audio output devices 410, 412, and 414 can operate with high amplitude within this low-frequency spectrum. Snare drum 440, tenor saxophone 462, snare drum 442, and French horn 454, which primarily generate output in the mid-frequency spectrum and less output in the low-frequency range, are assigned to audio output devices 420, 422, 424, and 426, respectively. Audio output devices 420, 422, 424, and 426 have frequency operating ranges within the mid-frequency spectrum and are unable to output high-amplitude sounds in the low-frequency range. Trumpet 450, trombone 452, vocal part / microphone 470, and hi-hat 446, which generate audio with a relatively high frequency spectrum, are assigned to audio output devices 430, 432, 434, and 436, respectively. Audio output devices 430, 432, 434, and 436 have frequency operating ranges within the high-frequency spectrum and lack the ability to play high-amplitude sounds in the low-frequency range.
[0057] Figure 5 Another example of a listening environment 500 for coordinating audio systems according to one or more aspects of various embodiments is shown. As shown, the listening environment 500 includes four audio output devices 510, 512, 520, and 530. The four audio output devices 510, 512, 520, and 530 play component audio signals corresponding to various instruments and voices. Audio output device 510 primarily plays component audio signals corresponding to drums 540. Audio output device 512 primarily plays component audio signals corresponding to a mix of bass guitar 552 and lead guitar 550. Audio output device 520 primarily plays component audio signals corresponding to a mix of keyboard 554 and two backing vocalists 572 and 574. Audio output device 530 primarily plays component audio signals corresponding to lead vocalist / microphone 570.
[0058] The listening environment 500 generated by audio output devices 510, 512, 520, and 530 provides different audio experiences depending on the user's location within the listening environment 500. As shown, user 580 is centrally located between audio output devices 510, 512, 520, and 530. Therefore, user 580 hears a balanced mix of component audio signals for drums 540, bass guitar 552, lead guitar 550, keyboard 554, lead vocals / microphone 570, and backing vocals / microphones 572 and 574. User 582 is located near audio output device 512. Therefore, user 582 primarily hears the component audio signals for bass guitar 552 and lead guitar 550, with each of the remaining component audio signals at a lower level. User 584 is located near audio output device 530. Therefore, user 584 primarily hears the component audio signal for lead vocals 570, with each of the remaining component audio signals at a lower level. User 586 is located near audio output device 510. Thus, user 586 primarily hears the component audio signal for drums 540, and to a lesser extent each of the remaining component audio signals. User 588 is located near audio output device 520. Thus, user 588 primarily hears the component audio signal for keyboard 554 and two backing vocalists 572 and 574, and to a lesser extent each of the remaining component audio signals. In one example, the component audio signal for keyboard 554 is played from audio output devices 520 and 530, and user 588 hears a phantom image of the keyboard at a location between audio output devices 520 and 530, even though there is no actual audio output device at that location.
[0059] In some examples, the audio upmixing application 152 assigns component audio signals to audio output devices based on matching the frequency spectra of the component audio signals with the frequency operating ranges of the respective audio output devices. In this regard, some audio output devices have frequency responses that prevent them from playing the lowest notes at the same output level as mid- and high-frequency notes. Consequently, these audio output devices are less suitable for reproducing signals from instruments with high-amplitude, low-frequency notes. As shown in the figure, drums 540, which generate audio with relatively high amplitude in the low-frequency spectrum, are assigned to audio output device 510. Audio output device 510 is capable of operating at high amplitude within this low-frequency spectrum. As shown in the figure, bass guitar 552 and lead guitar 550, which generate audio with relatively high amplitude in the low-frequency spectrum, are assigned to audio output device 512. Audio output device 512 is capable of operating at high amplitude within this low-frequency spectrum. Keyboard 554, which primarily generates output in the mid-frequency spectrum and less output in the low-frequency range, and two backing vocalists 572 and 574 are assigned to audio output device 520. The audio output device 520 has a frequency operating range within the mid-frequency spectrum and is unable to output high-amplitude sounds in the low-frequency range. The lead singer / microphone 570, which generates audio with a relatively high frequency spectrum, is assigned to the audio output device 530. The audio output device 530 has a frequency operating range within the high-frequency spectrum and lacks the ability to play high-amplitude sounds in the low-frequency range.
[0060] Figure 6 is a flowchart of method steps for generating a set of audio streams for a coordinated audio system according to one or more aspects of various embodiments. Figures 1 to 5 Although the systems and examples are described herein, those skilled in the art will understand that any system configured to perform the method steps in any order is within the scope of the various embodiments.
[0061] As shown, method 600 begins at step 602, where an audio upmixing application 152 executing on a computing device 180 locates audio output devices 126 in a listening environment. The audio upmixing application 152 locates audio output devices 126 that are currently reachable by the computing device 180 via one or more wired and / or wireless networks. In this regard, the audio upmixing application 152 includes audio output devices 126 that have recently entered a network accessible to the computing device 180. Similarly, the audio upmixing application 152 excludes audio output devices 126 that have recently exited a network accessible to the computing device 180. Thus, the audio upmixing application 152 adapts to the current set of audio output devices 126 as these audio output devices 126 enter, exit, and move within the listening environment.
[0062] At step 604, the audio upmixing application 152 receives one or more audio input signals. In some examples, the audio upmixing application 152 receives the audio input signals from the input device 122. In some examples, the audio upmixing application 152 retrieves data representing the audio input signals from the database 142 stored in the memory 116, from the storage device 114, or the like. Typically, the input audio signal is a stereo audio signal including a left audio channel and a right audio channel. Alternatively, the audio input signal is a mono audio signal having a single channel. In some examples, the audio input signal is a multi-channel encoded audio signal including two or more channels, such as four channels, six channels, eight channels, or the like. These multi-channel encoded audio signals include four-channel audio, DVD-Audio, Super Audio CD, Dolby Atmos, and the like.
[0063] At step 606, the audio upmixing application 152 separates a plurality of component audio signals from one or more audio input signals. In some embodiments, the audio upmixing application 152 performs a signal separation process on an audio input signal (such as a stereo audio signal) to separate the plurality of component audio signals from the audio input signal. In some examples, the audio input signal comprises a song performed by a rock band comprising a vocalist, a lead guitar, a bass guitar, and drums. The audio upmixing application 152 performs a signal separation process on the audio input signal to separate the audio input signal into four component audio signals, one audio signal for each of the vocalist, the lead guitar, the bass guitar, and the drums. In some examples, the audio input signal comprises a song performed by a ten-piece jazz band comprising various horns and drums as well as the vocalist. The audio upmixing application 152 performs a signal separation process on the audio input signal to separate the audio input signal into eleven component audio signals, one audio signal for the vocalist and one audio signal for each instrument in the ten-piece band. In some examples, the audio input signal includes a song performed by a pop band that includes a lead vocalist, two backing vocalists, a lead guitar, a keyboard, and drums. The audio upmixing application 152 performs a signal separation process on the audio input signal to separate the audio input signal into seven component audio signals, one audio signal for each of the lead vocalist, the two backing vocalists, the lead guitar, the bass guitar, the keyboard, and the drums.
[0064] At step 608, the audio upmixing application 152 maps the component audio signals included in the subset of component audio signals to one or more audio output devices 126 in the listening environment. The subset of component audio signals includes two or more component audio signals included in the plurality of component audio signals in step 606, up to and including all component audio signals. In some examples, the audio upmixing application 152 maps each component audio signal included in the subset of component audio signals to a different audio output device 126 in a one-to-one mapping. As a result, each audio output device plays a different component audio signal. In some examples, the audio upmixing application 152 maps two or more component audio signals to the same audio output device 126 in a two-to-one mapping or a many-to-one mapping. As a result, the audio output device 126 plays a mix of the two or more component audio signals. In some examples, the audio upmixing application 152 maps the component audio signals to two or more audio output devices 126 in a one-to-two mapping or a one-to-many mapping. As a result, multiple audio output devices 126 can play a portion of the same component audio signal. The audio upmixing application 152 may employ these mapping techniques in any technically feasible combination.In some examples, a certain component audio signal is omitted from playback and is not mapped to any audio output device 126, such as in a karaoke use case.
[0065] At step 610, the audio upmixing application 152 transmits the component audio signals included in the subset of component audio signals to corresponding audio output devices 126 based on the mapping determined in step 608. The audio output devices 126 output audio corresponding to the component audio signals from the audio input signal received from the audio upmixing application 152. In some embodiments, data corresponding to the component audio signals can be transmitted from the audio upmixing application 152 to at least one audio output device 126, and the data can be transmitted between audio output devices 126.
[0066] The method 600 then returns to step 602 to locate the audio output device 126 currently in the listening environment. In this way, the method 600 dynamically changes the upmix as audio output devices move within the listening environment, leave the network, new audio output devices enter the network, and so on.
[0067] In summary, a computing device and multiple audio output devices are coupled together to form a coordinated audio system. The computing device decomposes a received audio signal (such as a music signal) into separate audio streams based on certain criteria. The disclosed technology generates a customizable upmix composed of multiple audio streams in real time, each of which is transmitted to one or more audio output devices. In some examples, the computing device generates a different audio stream for each instrument and voice in the received audio signal. Each of these separate audio streams is then played on one or more audio output devices, thereby creating an immersive audio field between and within these audio output devices. As a user moves within the listening environment, they hear different customized mixes with varying balances of instruments and voices, depending on their relative distance from each of the audio output devices. In some examples, users can move to different locations to hear all instruments and voices in a balanced manner, or they can move to various locations where one or more separate audio streams predominate. As users move within this audio field, each user experiences a different mix of the separate audio streams, depending on their position and orientation within the audio field generated by the audio output devices. As a result, users can enjoy a more interactive and immersive listening experience compared to traditional technologies.
[0068] At least one technical advantage of the disclosed technology over the prior art is that it enables the deployment of multiple audio output devices (such as personal speakers) in party mode without generating the undesirable combing effect common with conventional techniques. Another technical advantage of the disclosed technology over the prior art is that each user within an environment can have a different audio experience based on their position relative to the multiple audio output devices. Furthermore, the upmixes transmitted to the multiple audio output devices dynamically change based on the speakers in the network, the source audio being used, and so on. More specifically, the technology dynamically changes the upmixes as audio output devices move within the environment, leave the network, or new audio output devices enter the network. In this way, the upmixes are adapted to generate an appropriate sound field as the audio output devices change over time. Consequently, users can enjoy a more interactive and immersive listening experience compared to conventional techniques. These technical advantages provide one or more technical improvements over prior art approaches.
[0069] 1. In some embodiments, a computer-implemented method for generating an audio signal in an audio system includes: receiving an audio input signal; separating a plurality of component audio signals from the audio input signal; for a subset of component audio signals included in the plurality of component audio signals, mapping each component audio signal included in the subset of component audio signals to one or more of a plurality of audio output devices; and transmitting each component audio signal to the corresponding one or more audio output devices based on the mapping.
[0070] 2. The computer-implemented method of clause 1, wherein each component audio signal included in the plurality of component audio signals is mapped to a different audio output device included in the plurality of audio output devices.
[0071] 3. A computer-implemented method according to clause 1 or clause 2, wherein a first component audio signal included in the plurality of component audio signals is mapped to a first audio output device included in the plurality of audio output devices, and a second component audio signal included in the plurality of component audio signals is mapped to the first audio output device.
[0072] 4. A computer-implemented method according to any of clauses 1 to 3, wherein a first component audio signal included in the plurality of component audio signals is mapped to a first audio output device included in the plurality of audio output devices, and the first component audio signal is also mapped to a second audio output device included in the plurality of audio output devices.
[0073] 5. The computer-implemented method of any of clauses 1 to 4, further comprising: receiving user input identifying a first component audio signal included in the plurality of component audio signals; and removing the first component audio signal from the plurality of component audio signals to create the subset of component audio signals.
[0074] 6. A computer-implemented method according to any of clauses 1 to 5, wherein each component audio signal included in the plurality of component audio signals comprises a different instrument, sound or group of instruments or sounds included in the audio input signal.
[0075] 7. A computer-implemented method according to any one of clauses 1 to 6, wherein the plurality of audio output devices are connected via a network, and the method further comprises: determining that a first audio output device not included in the plurality of audio output devices is connected to the network; adding the first audio output device to the plurality of audio output devices to create an updated plurality of audio output devices; for each component audio signal included in the plurality of component audio signals, mapping the component audio signal to one or more audio output devices included in the updated plurality of audio output devices; and transmitting each component audio signal to the corresponding one or more audio output devices based on the mapping.
[0076] 8. A computer-implemented method according to any one of clauses 1 to 7, wherein the plurality of audio output devices are connected via a network, and the method further comprises: determining that a first audio output device included in the plurality of audio output devices is no longer connected to the network; removing the first audio output device from the plurality of audio output devices to create an updated plurality of audio output devices; for each component audio signal included in the plurality of component audio signals, mapping the component audio signal to one or more of the updated plurality of audio output devices; and transmitting each component audio signal to the corresponding one or more audio output devices based on the mapping.
[0077] 9. The computer-implemented method of any of clauses 1 to 8, wherein the mapping is based on at least one of a spectrum or a volume of a first component audio signal included in the plurality of component audio signals.
[0078] 10. A computer-implemented method according to any one of clauses 1 to 9, wherein the mapping is based on a virtual position of a first component audio signal included in the plurality of component audio signals, wherein the virtual position is determined when the plurality of component audio signals are mixed and mastered.
[0079] 11. In some embodiments, one or more non-transitory computer-readable storage media include instructions that, when executed by one or more processors at a first computing device, cause the one or more processors to perform the following steps: receiving an audio input signal; separating a plurality of component audio signals from the audio input signal; for a subset of component audio signals included in the plurality of component audio signals, mapping each component audio signal included in the subset of component audio signals to one or more of a plurality of audio output devices; and transmitting each component audio signal to the corresponding one or more audio output devices based on the mapping.
[0080] 12. The one or more non-transitory computer-readable media of clause 11, wherein the steps further comprise: receiving user input identifying a first component audio signal included in the plurality of component audio signals; and removing the first component audio signal from the plurality of component audio signals to create the subset of component audio signals.
[0081] 13. One or more non-transitory computer-readable storage media according to clause 11 or clause 12, wherein each component audio signal included in the plurality of component audio signals comprises a different instrument, sound, or group of instruments or sounds included in the audio input signal.
[0082] 14. One or more non-transitory computer-readable storage media according to any one of clauses 11 to 13, wherein the plurality of audio output devices are connected via a network, and wherein the steps further comprise: determining that a first audio output device not included in the plurality of audio output devices is connected to the network; adding the first audio output device to the plurality of audio output devices to create an updated plurality of audio output devices; for each component audio signal included in the plurality of component audio signals, mapping the component audio signal to one or more audio output devices included in the updated plurality of audio output devices; and transmitting each component audio signal to the corresponding one or more audio output devices based on the mapping.
[0083] 15. One or more non-transitory computer-readable storage media according to any one of clauses 11 to 14, wherein the plurality of audio output devices are connected via a network, and wherein the steps further comprise: determining that a first audio output device included in the plurality of audio output devices is no longer connected to the network; removing the first audio output device from the plurality of audio output devices to create an updated plurality of audio output devices; for each component audio signal included in the plurality of component audio signals, mapping the component audio signal to one or more of the updated plurality of audio output devices; and transmitting each component audio signal to the corresponding one or more audio output devices based on the mapping.
[0084] 16. One or more non-transitory computer-readable storage media according to any of clauses 11 to 15, wherein the mapping is based on at least one of a spectrum or a volume of a first component audio signal included in the plurality of component audio signals.
[0085] 17. The one or more non-transitory computer-readable storage media of any of clauses 11 to 16, wherein mapping the component audio signals to one or more of a plurality of audio output devices is based on input received from a user interface.
[0086] 18. One or more non-transitory computer-readable storage media as described in any of clauses 11 to 17, wherein the steps further include: receiving a second audio input signal; separating a second plurality of component audio signals from the second audio input signal; for each second component audio signal included in the second plurality of component audio signals, mapping the component audio signal to one or more of the plurality of audio output devices; and transmitting each second component audio signal to the corresponding one or more audio output devices based on the mapping.
[0087] 19. In some embodiments, a computing device comprises: a memory storing an application; and one or more processors, wherein the one or more processors are configured, when executing the application, to: receive an audio input signal; separate a plurality of component audio signals from the audio input signal; for a subset of component audio signals included in the plurality of component audio signals, map each component audio signal included in the subset of component audio signals to one or more of a plurality of audio output devices; and transmit each component audio signal to the corresponding one or more audio output devices based on the mapping.
[0088] 20. The computing device of clause 19, wherein the computing device is coupled to the plurality of audio output devices via a wireless network.
[0089] Any and all combinations of any claim elements recited in any claim and / or any elements described in this application, in any manner, are within the intended scope of this disclosure and protection.
[0090] The description of the various embodiments has been presented for purposes of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments.
[0091] Aspects of the present embodiment may be embodied as a system, method or computer program product. Therefore, aspects of the present disclosure may take the following forms: a complete hardware embodiment, a complete software embodiment (including firmware, resident software, microcode, etc.) or an embodiment combining software aspects and hardware aspects, which embodiments may generally be referred to as "modules", "systems" or "computers" herein. In addition, any hardware and / or software technology, process, function, component, engine, module or system described in the present disclosure may be implemented as a collection of circuits or circuits. In addition, aspects of the present disclosure may be in the form of a computer program product embodied in one or more computer-readable media, on which computer-readable program code is embodied.
[0092] Any combination of one or more computer-readable media can be utilized. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media will include the following media: an electrical connection with one or more wires, a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium can be any tangible medium that can contain or store a program for use by an instruction execution system, device, or apparatus, or that is connected to an instruction execution system, device, or apparatus.
[0093] Aspects of the present disclosure are described above with reference to the flowchart illustrations and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart illustration and / or block diagram and the combination of the boxes in the flowchart illustration and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine. These instructions enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram when executed via the processor of a computer or other programmable data processing device. Such processors may be, but are not limited to, general-purpose processors, special-purpose processors, special-purpose processors or field programmable gate arrays.
[0094] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, segment, or portion of a code that includes one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions mentioned in the boxes may not appear in the order mentioned in the accompanying drawings. For example, depending on the functionality involved, two boxes shown in succession may be executed substantially simultaneously, or the boxes may sometimes be executed in reverse order. It should also be noted that each box in the block diagram and / or flowchart illustration, and the combination of boxes in the block diagram and / or flowchart illustration, may be implemented by a special-purpose hardware-based system that performs the specified function or action, or by a combination of special-purpose hardware and computer instructions.
[0095] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, which is determined by the claims that follow.
Claims
1. A computer-implemented method for generating an audio signal in an audio system, the method comprising: receiving an audio input signal; separating a plurality of component audio signals from the audio input signal; For a subset of component audio signals included in the plurality of component audio signals, mapping each component audio signal included in the subset of component audio signals to one or more of a plurality of audio output devices; and Each component audio signal is transmitted to a corresponding one or more audio output devices based on the mapping. 2 . The computer-implemented method of claim 1 , wherein each component audio signal included in the plurality of component audio signals is mapped to a different audio output device included in the plurality of audio output devices.
3. The computer-implemented method of claim 1 , wherein a first component audio signal included in the plurality of component audio signals is mapped to a first audio output device included in the plurality of audio output devices, and a second component audio signal included in the plurality of component audio signals is mapped to the first audio output device.
4. The computer-implemented method of claim 1 , wherein a first component audio signal included in the plurality of component audio signals is mapped to a first audio output device included in the plurality of audio output devices, and the first component audio signal is also mapped to a second audio output device included in the plurality of audio output devices.
5. The computer-implemented method of claim 1 , further comprising: receiving user input identifying a first component audio signal included in the plurality of component audio signals; as well as The first component audio signal is removed from the plurality of component audio signals to create the subset of component audio signals.
6. The computer-implemented method of claim 1, wherein each component audio signal included in the plurality of component audio signals comprises a different instrument, sound, or group of instruments or sounds included in the audio input signal.
7. The computer-implemented method of claim 1 , wherein the plurality of audio output devices are connected via a network, and the method further comprising: determining that a first audio output device not included in the plurality of audio output devices is connected to the network; adding the first audio output device to the plurality of audio output devices to create an updated plurality of audio output devices; for each component audio signal included in the plurality of component audio signals, mapping the component audio signal to one or more audio output devices in the updated plurality of audio output devices; as well as Each component audio signal is transmitted to a corresponding one or more audio output devices based on the mapping.
8. The computer-implemented method of claim 1 , wherein the plurality of audio output devices are connected via a network, and the method further comprising: determining that a first audio output device included in the plurality of audio output devices is no longer connected to the network; removing the first audio output device from the plurality of audio output devices to create an updated plurality of audio output devices; for each component audio signal included in the plurality of component audio signals, mapping the component audio signal to one or more of the updated plurality of audio output devices; as well as Each component audio signal is transmitted to a corresponding one or more audio output devices based on the mapping.
9. The computer-implemented method of claim 1, wherein the mapping is based on at least one of a spectrum or a volume of a first component audio signal included in the plurality of component audio signals.
10. The computer-implemented method of claim 1, wherein the mapping is based on a virtual position of a first component audio signal included in the plurality of component audio signals, wherein the virtual position is determined when the plurality of component audio signals are mixed and mastered.
11. One or more non-transitory computer-readable storage media comprising instructions that, when executed by one or more processors at a first computing device, cause the one or more processors to perform the following steps: receiving an audio input signal; separating a plurality of component audio signals from the audio input signal; For a subset of component audio signals included in the plurality of component audio signals, mapping each component audio signal included in the subset of component audio signals to one or more of a plurality of audio output devices; and Each component audio signal is transmitted to a corresponding one or more audio output devices based on the mapping.
12. The one or more non-transitory computer-readable storage media of claim 11, wherein the steps further comprise: receiving user input identifying a first component audio signal included in the plurality of component audio signals; as well as The first component audio signal is removed from the plurality of component audio signals to create the subset of component audio signals.
13. The one or more non-transitory computer-readable storage media of claim 11, wherein each component audio signal included in the plurality of component audio signals comprises a different instrument, sound, or group of instruments or sounds included in the audio input signal.
14. The one or more non-transitory computer-readable storage media of claim 11, wherein the plurality of audio output devices are connected via a network, and wherein the steps further comprise: determining that a first audio output device not included in the plurality of audio output devices is connected to the network; adding the first audio output device to the plurality of audio output devices to create an updated plurality of audio output devices; for each component audio signal included in the plurality of component audio signals, mapping the component audio signal to one or more audio output devices in the updated plurality of audio output devices; as well as Each component audio signal is transmitted to a corresponding one or more audio output devices based on the mapping.
15. The one or more non-transitory computer-readable storage media of claim 11, wherein the plurality of audio output devices are connected via a network, and wherein the steps further comprise: determining that a first audio output device included in the plurality of audio output devices is no longer connected to the network; removing the first audio output device from the plurality of audio output devices to create an updated plurality of audio output devices; for each component audio signal included in the plurality of component audio signals, mapping the component audio signal to one or more of the updated plurality of audio output devices; as well as Each component audio signal is transmitted to a corresponding one or more audio output devices based on the mapping. 16 . The one or more non-transitory computer-readable storage media of claim 11 , wherein the mapping is based on at least one of a spectrum or a volume of a first component audio signal included in the plurality of component audio signals.
17. The one or more non-transitory computer-readable storage media of claim 11, wherein mapping the component audio signals to one or more of a plurality of audio output devices is based on input received from a user interface.
18. The one or more non-transitory computer-readable storage media of claim 11, wherein the steps further comprise: receiving a second audio input signal; separating a second plurality of component audio signals from the second audio input signal; for each second component audio signal included in the second plurality of component audio signals, mapping the component audio signal to one or more of the plurality of audio output devices; as well as Each second component audio signal is transmitted to a corresponding one or more audio output devices based on the mapping.
19. A computing device comprising: a memory storing an application program; as well as One or more processors, wherein the one or more processors, when executing the application, are configured to: Receive audio input signal, separating a plurality of component audio signals from the audio input signal, For a subset of component audio signals included in the plurality of component audio signals, mapping each component audio signal included in the subset of component audio signals to one or more of a plurality of audio output devices, and Each component audio signal is transmitted to a corresponding one or more audio output devices based on the mapping.
20. The computing device of claim 19, wherein the computing device is coupled to the plurality of audio output devices via a wireless network.