System and method for immersive music performance between at least two remote locations on a network

The binauralization technique addresses the challenge of immersive remote music performance by creating a virtual acoustic space with scalable zone mapping and low-latency audio streaming, ensuring realistic sound localization and reduced complexity for networked music systems.

JP2026512900APending Publication Date: 2026-04-21BONZA MUSIC LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
BONZA MUSIC LTD
Filing Date
2024-03-26
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing network-based music performance systems fail to provide a realistic immersive experience for remote musicians due to localized sound sources within headphones, leading to unnatural auditory-visual inconsistencies and increased computational complexity with the number of participants.

Method used

A binauralization technique using acoustic measurements to create a virtual acoustic space, allowing remote participants to perceive sound sources outside the head, with scalable zone mapping and low-latency audio streaming to maintain immersion and minimize network bandwidth.

Benefits of technology

Enables realistic musical interaction and immersion for multiple participants by accurately localizing sound sources within a shared virtual space, reducing computational and network complexity, and minimizing latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026512900000001_ABST
    Figure 2026512900000001_ABST
Patent Text Reader

Abstract

A system and method for immersive musical performance between at least two remote locations on a network. A system and method for concerto musical performance in which performers in a first location space and a second location space located away from the first location space can experience the perception that they are in the same location. The method / system requires acquiring at least one binaural room impulse response of a desired space (which may or may not be one of those locations), transmitting a low-latency audio stream of the performance in each location space over the network, and applying the binaural room impulse response (BRIR) as a real-time filter. In this way, sound sources from remote locations are perceived as being in the desired space when played back through headphones. In one embodiment, one or both of the location spaces can be divided into zones, and different BRIRs are applied depending on the location of the zone corresponding to the location where the BRIR was measured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0005]

[0001] The present invention relates to a system and method for facilitating a music performance between at least two remote locations on a network, for example, a system and method for creating an immersive virtual acoustic space in which multiple players can play together while remotely connected over the Internet.

Background Art

[0002] Low-latency network-based software solutions for simultaneous performance at two or more remote locations are known. Specific examples are Soundjack (trademark) and SonoBus (trademark), which make it possible to deliver a standard stereo audio experience via headphones.

[0003] The problem that becomes apparent in such a system is that the participants do not have the perception of playing together in the same space, that is, local musicians do not feel as if external musicians are in the same room. This problem is caused by the fact that all sound sources are localized "inside the head" by headphone listening, and the sound comes to be heard unnaturally and feels unnatural. Especially when presented on a screen, there is no consistency between the auditory cue and the visual cue. <​​​​​For a performer, the perceived relative position of other performers is more important than any delay between the sound event and its visual equivalent, such as when the video of a drumstick hitting a cymbal or snare drum does not match the sound.

[0006] An example of a remote performance system known in this field is International Publication 2022 / 196073, which relies on measuring binaural room impulse responses (BRIR) from the positions of each performer in space. U.S. Patent Application 2022 / 0114993 discloses a gesture-controlled virtual instrument, but does not discuss spatial audio, binaural sound, or immersive audio. International Publication 2018 / 116368 discloses an HRTF framework for directly rendering performer positions. Generally, the computational complexity of prior art systems increases significantly with the number of participants. [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] This invention seeks to address the above-mentioned problems revealed in systems available for playing music together over a network. At the very least, this invention provides an alternative experience for users who wish to play music together over the internet (for example, as practice or in a formal performance). This invention can be applied particularly to music education. [Means for solving the problem]

[0008] Broad aspects of the present invention will be outlined in accordance with claim 1 of the attached claims.

[0009] This invention utilizes a binauralization technique to create the impression for the user that the sound source is outside the head when listening with headphones. The invention is implemented by performing acoustic measurements in an actual room to capture the acoustic characteristics of the space, and using these characteristics as a real-time filter so that the sound source is perceived as being within that space. The present method of the invention described in the claims differs significantly from prior art, such as that described in International Publication No. 2022 / 196073, because the position of a remote participant is not directly binauralized, but rather presented on a virtual representation of the loudspeaker system. This offers a significant advantage in that the end user can define the performer's position anywhere within the loudspeaker's sound field, as desired.

[0010] Binaural processing is performed for multiple listeners to create an immersive virtual space. In some embodiments, binaural measurements are pre-processed before auditory processing, i.e., before the procedure designed to model and simulate the experience of acoustic phenomena rendered as a sound field in a virtual space. Without this processing, the preceding sound effect would cause strong localization errors, leading to an unrealistic experience.

[0011] The method of the present invention is scalable to any number of participants. For example, zone mapping depends on the geometric conditions of the reproducible space rather than the number of listeners / participants. In some forms, the method can be configured to capture acoustics from both spaces (e.g., the external space in a stereo mix and the local space in binaural rendering), or to exclusively capture only the local space or only the external space. According to an aspect of the concept of the present invention, as described, for example, in claim 15, the method utilized herein includes obtaining a virtual loudspeaker configuration of a desired space (e.g., a first positional space where performers are positioned and / or another space, e.g., a famous venue) to provide the acoustic characteristics of the aforementioned desired space. The virtual loudspeaker configuration is obtained by performing binaural acoustic measurements at the center of at least one zone within the desired space for at least two loudspeakers. Preferably, there are three speakers, e.g., a left speaker, a center speaker, and a right speaker. In this way, the virtual loudspeaker configuration can be applied as a real-time filter so that when played back through headphones, the sound source is perceived as being within the desired space.

[0012] The present invention has several advantages, namely, • Realistic musical interaction resulting from the sharing of a virtual space and the immersion experienced by participants within that virtual space. This system is particularly applicable to network music education applications, but is not limited to this use. • The spatial separation and localization of sound sources are designed to allow participants to easily focus on individual performers. • Using a "zone" approach, it can be scaled to accommodate any number of local / remote performers without increasing the complexity of local rendering. • By setting a minimum of two channels as the bandwidth requirement "per site" (for example, assuming 1 Mbps per channel per site, resulting in approximately 800 kbps at 48 kHz, 16 bits, with approximately 200 kbps reserved for other network processes such as technical communications using 64 kbps opus compression), it is possible to minimize the required network bandwidth while enabling computationally efficient indoor simulation and binaural delivery. • By counting network latency as the initial time gap in auditory perception, air propagation delay is eliminated from the immersive audio NMP (Networked Music Performance) experience. This could lead to... [Brief explanation of the drawing]

[0013] The present invention will be described with reference to the accompanying drawings. [Figure 1] A conceptual plan view of the present invention is shown. [Figure 2] This shows a conceptual plan of the implementation method in the reproduction framework. [Figure 3] This document presents an acoustic measurement framework. [Figure 4] This shows a conceptual plan of the implementation method in the multi-site reproduction framework. [Figure 5] A diagram illustrating the hardware components is provided. [Figure 6] This shows an example of audio routing at a single location. [Figure 7] This shows an overview of the audio processing signal flow. [Figure 8] This outlines the UDP audio streaming method. [Figure 9] This demonstrates a convolution process using the Fast Fourier Transform (FFT) superposition method. [Modes for carrying out the invention]

[0014] The following description shows exemplary embodiments and helps to explain the principles of the present invention together with the drawings. However, since modifications will be apparent to those skilled in the art and are considered to be covered by this specification, the scope of the present invention is not intended to be limited to the exact details of the embodiments. The terms of components used in this specification shall be given a broad interpretation that also includes equivalent functions and features. In some cases, several alternative terms (synonyms) for a feature are provided, but these terms are not intended to be exhaustive.

[0015] Explanatory terms should also be given the broadest possible interpretation. For example, when interpreting each description in this specification that uses the term "comprising", the term "comprising" should be interpreted to mean "consisting at least in part of" so that other features may exist or features other than those preceded by that term may exist. Related terms such as "comprise" and "comprises" should be interpreted in the same way. Directional terms such as "vertical", "horizontal", "upper", "lower", "above", and "below" are usually used for the convenience of explanation with reference to the drawings and are not intended to be ultimately limiting if equivalent functions can be achieved with alternative dimensions and / or directions.

[0016] The description in this specification refers to embodiments having a particular combination of steps or features, but it is considered that further combinations and hybrid combinations of compatible steps or features between embodiments are possible. In fact, an isolated feature can function as an invention independently of other features and does not necessarily have to be implemented as a complete combination.

[0017] It will be understood that the illustrated embodiments are shown for illustrative purposes only. In practice, the present invention can be applied to many different configurations that are easy for those skilled in the art to implement.

[0018] The proposed solution utilizes a novel framework that uses the distribution of binaural measurement points around a virtual loudspeaker configuration to facilitate the capture and reproduction of acoustic spaces. This solution is exemplified, among other things, by two measurement and reproduction paradigms, such as the "the other person is here" paradigm where an external participant in a musical experience is virtually present within the local space, or the "you are on the other side" paradigm where a shared virtual environment is presented to all parties involved in a networked musical experience. The present invention is best explained by reference to the drawings and the above concepts.

[0019] FIG. 1 represents a schematic implementation form of the present invention. For example, a conceptual layout 10 of performers is shown, and a group of local performers 11 to 16 located in a first position space are arranged with respect to a display device, such as a screen 17. In a second position space away from the first position space, one or more remote performers 18 are also facing the display device. In this conceptual layout 10, the screen 17 is represented as the same device, but in reality, display devices exist at both ends of the network connection, functioning as a "window" between the performers, like a recording studio having separate rooms for separating performances.

[0020] Next, referring to FIG. 2, according to the "the other person is here" implementation method in the reproduction framework, an active local listening area is divided into a plurality of listening zones, such as three zones Z1, Z2, and Z3 (also known as zones A, B, and C), each accommodating a pair of local performers, such as 11 and 12, 13 and 14, and 15 and 16. However, as long as the physical space permits, additional performers may also participate in the zones. The screen area 17 is acoustically represented by three virtual loudspeakers, such as left 19, center 20, and right 21. The measurement of these virtual loudspeakers is achieved by binaural acoustic measurements from each loudspeaker to the center of each listening zone Z1, Z2, and Z3 (as explained by referring to FIG. 3), and finally reproduced by the user within headphones. For example, the performer 18 looking at the group of remote performers through the "window" of the screen 17 hears a stereo representation based on which zone performers 11 to 16 are located in.

[0021] The performer's two-channel stereo mix is ​​received at each local reproduction site along with the video signal, although the video signal is not essential to achieving the improved auditory experience of the present invention. However, the presence of the video signal allows the spatial position of the remote performer in the second spatial space to coincide with a visual cue, i.e., where the performer is standing. Simultaneous and equivalent transmission takes place from the local site to the remote site.

[0022] While the stereo mix is ​​transmitted, reproduction is performed in 3 channels, and if the reproduction angle becomes too large, for example, if the left and right virtual loudspeakers exceed ±45 degrees depending on the geometric conditions of the room, center imaging is enhanced. The center channel is extracted by adding the L+R channels and mixed according to preference (typically -6dB for large screen widths, for example). If the reproduction angle is less than ±30 degrees, the center channel is not necessary, and the system can be simplified to only 2 channels.

[0023] As shown in Figure 2, all performers 11 through 16 and 18 wear headphones, binaural presentation is rendered in real time, and performers perceive the position of auditory events so that they come from the correct position corresponding to the musician on the screen. For example, musician 18, who is in a distant location, hears the performance of performers 15 and 16, who are positioned in the stereo sound field on the left, because from musician 18's perspective, performers 15 and 16, who are in zone Z3, appear to be on the left side of the screen. The simplest embodiment of the present invention is when all performers are stationary, but in some embodiments, performers can move, and their movement on the screen is tracked and rendered in the stereo sound field so that the headphone mix for the local performer is panned in the direction of their movement. In other words, a further embodiment of the present invention tracks a remote performer by motion detection means, and mix adjustments can be made in real time, for example, when a performer crosses the stage. In this way, a remote performer moving from a left position to a right position can have the mix of their instruments panned from the far left to the center.

[0024] In an exemplary form, binaural acoustic measurements are processed so that correct stereo imaging is perceived by a participant, e.g., 18, for each listening zone Z1, Z2, and Z3. Thus, local performers 11 and 12, standing on the left side in front of screen 17, are perceived by remote performer 18 as being on the right side compared to the central mix of performers 13 and 14.

[0025] The processing according to the present invention involves manipulating binaural cues of time and level differences between the two ears. Without such processing, the precedence effect can cause localization errors in the direction of the nearest virtual loudspeaker. The precedence effect, or the law of the first wavefront, is a binaural psychoacoustic effect in which a listener perceives a single auditory event when one sound is followed by another sound separated by a sufficiently short time delay (below the listener's echo threshold). That is, the perceived spatial position is governed by the position of the first sound to arrive (the first wavefront), and the delayed sound also affects the perceived position. However, this effect is suppressed by the precedence effect.

[0026] For optimal use, the performer should be in a fixed position, i.e., the instrument should be positioned facing the screen. For example, a pianist sitting in front of a grand piano who must move their head sideways to look at the screen and then move their head forward to look at the keyboard and play will have a suboptimal experience because the headphone mix may not account for head rotation. In this case, the piano keyboard should be parallel to the screen. However, in a further embodiment of the present invention, local ambisonics rendering can be implemented that utilizes omnidirectional binaural measurements, which can compensate for any arbitrary head rotation / movement.

[0027] Each listening zone Z1, Z2, and Z3 has an acoustic sweet spot at the measurement point; that is, the experience is most realistic for the listener 18 when the performers 11-16 are positioned within the zone corresponding to where the measurement was taken. However, the ventriloquism effect is strongly maintained within each zone, ensuring good localization on the screen. Outside the zones, the ventriloquism effect is weakened. The ventriloquism effect is an example of visual cues taking precedence over other senses such as hearing. In other words, the stereo image does not need to be perfect for the user to realistically perceive sound as coming from a specific direction, as long as it roughly matches the visual cues on the display 17.

[0028] The implementation described above is scalable to any number of participants. Zone mapping depends on the geometric conditions of the reproducible space, not on the number of listeners.

[0029] In one form, this method can incorporate the acoustics of both spaces, for example, the remote space in a stereo mix and the local space in binaural rendering, or (for example, when a remote instrument is picked up with a close-range microphone) the acoustics of the local space only.

[0030] In one application, a single binaural diffuse field measurement is also applied to / used for local performer monitoring of the performer's own instrument, so that the performer has the impression that room reverb is applied to their instrument, which is picked up by a close-range microphone (e.g., a clip-on microphone pointed into the bell of a trumpet). Another example is when an electronic piano keyboard is directly connected to a computer interface, i.e., when listening with headphones, there is usually no room acoustics, but diffuse field reverb provides the desired ambience.

[0031] In the second form of this method, which we refer to above as "you are on the other side," the same technique is employed, but the acoustic measurements used within the zone are not performed in the actual reproduced environment (i.e., the first spatial location), but instead, or in addition, in any desired acoustic environment under the correct relative geometric conditions. Such an environment could be the remote environment of another participant, or a completely different acoustic environment such as a famous recording studio or venue.

[0032] Figure 3 illustrates an exemplary method of the acoustic measurement framework, in which three binaural measurement positions 22, 23, and 24 corresponding to zones Z1, Z2, and Z3 are established relative to the actual loudspeaker positions 25, 26, and 27 on a line 28 representing the screen position.

[0033] Furthermore, diffuse sound field measurement and capture 29 can be performed at a position two to three times the critical distance point. In particular, Figure 3 is not to scale, and therefore position 29 may be much further back in the room than is shown. Direct sound is attenuated by the acoustic baffle 30 so that the measurement performed at position 29 provides a relatively neutral representation of the room reverb. As mentioned above, this neutral reverb is applied to the performer's own instrument (which may be a monaural signal) and can be mixed into the center of the performer's personal monitor mix.

[0034] As background, the Head-Related Transfer Function (HRTF) convolution process requires acoustic measurements of the binaural room impulse response (BRIR) in the space that needs to be simulated. This can be achieved using a KU100™ or similar binaural dummy head to perform impulse response measurements in the room to be simulated. Alternatively, ambient microphones can be used to capture ambisonic impulse response measurements and convert them into a binaural representation.

[0035] Following known methods for obtaining the acoustic characteristics of a room, three measurements are taken at positions 22, 23, and 24 using a binaural measurement device, i.e., a dummy head with stereo microphones. For example, sine sweep tones are emitted from speakers 25, 26, and 27 to sequentially excite the air for approximately 20 seconds each in the range of 20 Hz to 20 kHz, which is the range of human hearing. The length of the measurements typically depends on the reverb characteristics of the room. The output is saved as a stereo file for processing to obtain a deconvolved binaural room impulse response. In this example, each position measures the response from three speaker positions (i.e., a total of nine measurements). However, in a narrow sound field, only two speakers (without a center speaker) may be used. In a wider sound field, there may be four or more measurement positions corresponding to zones (Z1, Z2, Z3, etc.).

[0036] During measurement, standing and sitting positions should be considered (often depending on the type of instrument). As mentioned above, measurements are taken at each binaural position 22, 23, and 24 for each distance signal / tone emitted from each speaker 25, 26, and 27 (i.e., 3x3 measurements).

[0037] The specific measurement protocol described herein considers a virtual left-center-right (LCR) loudspeaker configuration as the sound source. There are three receiver positions 22, 23, and 24 defined with respect to the "zone" technique, and the binaural dummy head is positioned facing the center speaker, i.e., angled so that when at the outer edge positions 22 and 24, its eyes face the center speaker 26.

[0038] The measured BRIR is saved as a .wav file and can be used in the convolution process for binaural room simulation. The BRIR is applied locally in real time to the input low-latency audio signal of the remote performance.

[0039] If three or more different sites are used within the system, the replayed scene can be appropriately divided to match the split-screen visuals or multiple display screens. The screen that acts as a "window" to the remote performer can be configured to reproduce the audio panning of the relevant audio stream. The multi-site replay framework is illustrated in Figure 4, for example, in which the first and second groups of performers located at sites R1 and R2, respectively, are spatially positioned relative to the performer 31 at the third site behind the display screen 17.

[0040] It is worth noting that each group of performers can also have a corresponding headphone mix based on the spatial positioning of the remote performers on their local screen. For example, a performer in R2 can see group R1 on the left and a single performer 31 on the right on their screen, and the sound field in their headphone mix is ​​mixed accordingly. The system simply needs to track the relative positions on the location map in order to apply the correct BRIR stored locally to the input audio stream. To minimize audio stream latency, processing is performed locally.

[0041] The above description conceptually outlines the present invention, namely a system and method for delivering acoustic room simulation and binaural audio for remote immersive music networks. The system is designed for use over the internet between multiple remote locations, each with at least one musician, such as a band or orchestra, for applications such as music education where an instructor may be located remotely from one or more students. The system is expandable by incorporating a "zone" approach.

[0042] The system described herein utilizes low-latency audio streaming and rendering methods and may be limited by the capabilities of the public network. For example, to achieve the best results and realistic performance conditions, a maximum distance of approximately 500 km between locations (or 1000 km round trip between sites) is expected to be practical. However, since a single stereo mix is ​​exported from each location after the initial parameter setup, the distance is not affected by the number of users at a particular location. In any case, the distance limit may increase with improvements in communication technology.

[0043] The peer-to-peer streaming methods of the type described herein require bandwidth consumption that increases or decreases with the number of audio channels. A preferred embodiment involves two (or four, including the technical channel) streams between each pair of remote sites, allowing multiple sites to connect within reasonable bandwidth consumption. This is particularly beneficial when each site has multiple musicians, ensuring that bandwidth consumption does not increase even if musicians are added at a site.

[0044] In particular, bandwidth consumption and required network processing power do not increase with the size of the local group. Instead, the audio engine is designed to provide an immersive audio experience using the aggregated stereo image of all remote sites received from the streaming component. This approach is unique in the context of network music in that it does not require object-based audio (individual channels that make bandwidth management impossible) between sites, but still delivers binaural playback and binaural room simulation. In this way, immersive audio and room simulation are achieved within the constraints of low-bandwidth streaming.

[0045] Furthermore, the audio engine can be programmed in a way that requires only one instance of the audio rendering process, as opposed to generating a new audio rendering process for each individual remote performance group. This novel approach allows for controlling processing requirements within a reasonable level achievable, for example, on a home computer or embedded device.

[0046] The zone-based approach to rendering binaural audio is also noteworthy for enabling accurate localization of performers without the need to render separate mixes for each individual musician. By leveraging a ventriloquist effect, performers experience precise directional sound from visual cues of the other musician displayed on screen. This ensures that no hardware or digital routing and processing changes are required when new musicians join the group in each site / zone. This also avoids the increasing audio processing load that comes with group size.

[0047] In this context, an audio rendering method (e.g., BRIR convolution) is necessary because the split-superposition-addition method allows immersive audio processing to be achieved with minimal additional latency, and minimizing latency is crucial in the context of immersive audio. Essentially, this system introduces binaural immersive audio room simulation to the performance experience, thereby improving the aforementioned experience by simulating the experience of performing "in the room" with musicians located at a distance, and making networked musical interactions more natural. It is also worth noting that rendering with headphones rather than loudspeakers reduces latency due to sound propagation in the air (i.e., the speed of sound).

[0048] Figure 5 shows an overview of exemplary hardware components at each location, for example, including a first (optional) computer 32 used to record the local group's performance, a second computer 33 used for audio rendering and networking processes, an audio interface 34 used to provide audio input and output from the second computer 33, and a mixing console, mixer 35 for receiving input from a microphone 36 or DI capturing the local performance group. The mixer 35 sends the stereo mix to the second computer 33 via the audio interface input, along with additional auxiliary sends. In the exemplary configuration, mixing / panning is performed from a camera perspective, corresponding to what the remote user sees on their display screen.

[0049] One or more headphone amplifiers 37 may be provided, each receiving a binaural zone mix from an audio interface 34 for playback through headphones 38, one for each member of the local performance group. Locally, each performer can receive an individual monitor mix, including panning of fellow local musicians depending on their relative position. These require additional audio signals, but all are done locally and not streamed over a communication network. Otherwise, each local performer receives the same stereo mix of the remote components in their own monitor mix, as each performer generally sees the same image on the display screen in front of them.

[0050] In addition to the illustrated equipment, as mentioned, a video camera can capture images from a selected viewpoint (determining panning) for low-latency streaming by computer 33 or a separate computer. The video feed can be integrated with this system or run independently, i.e., using available video conferencing platforms such as Zoom®, Skype®, or Teams®.

[0051] Examples of equipment for single-location use are outlined in Table 1 below: [Table 1]

[0052] Examples of software components may include the JACK® audio connection kit used for routing audio between applications, digital audio workstations (DAWs) such as Reaper® used for hosting audio processing, convolution plugins such as the X-MCFX® convolver used to provide real-time convolution capabilities in DAWs, and Soundjack® used to provide low-latency audio streaming capabilities over communication networks.

[0053] Figure 6 shows an example of audio routing at a given location. First, the audio interface input (receiving a stereo mix of local performance from mixer 35, captured by, for example, microphone 36) is routed directly to a low-latency audio stream by audio router software and sent to the remote location, while the input low-latency audio stream from the remote location is routed to the DAW to apply a convolutional plugin based on the modeled space. In the exemplary form, each zone in the map of the virtual space has a separate set of binaural IRs, so the input stereo feed is rendered with the corresponding binaural IRs for each zone. The processed audio output is routed to the audio interface output and beyond for distribution to the performer at that location, for example, via headphone amplifier 37 and headphones 38. Depending on the zone in which the performer is located, there may be multiple stereo monitor mixes.

[0054] Figure 7 shows an overview of the audio processing signal flow. For example, a stereo mix representing a local performance group is received at block 40 from a desk (35). This is then transported to a remote site for auditory processing by, for example, a Soundjack® 41.

[0055] A stereo mix representing the composite stereo image of the remote performance group on screen is received at the site. A mono monitor mix 42 is also received from the desk for each zone (i.e., a total of three). This is added to the corresponding zone-specific binaural playback signal 43, which represents a virtual "stage wedge" monitor.

[0056] A mono diffuse reverb send 44 is received from the desk, and this send provides a binaural simulation of local performance groups, which is added to each zone mix for headphone playback 43.

[0057] Two private communication channels 45 are received from the audio interface input and can be passed to a streaming component for utility purposes. The audio engineer at the desk can speak to the remote site through the main mix. Microphone signals can also be input to the audio interface to add a "punch-in" option to the local audio mix.

[0058] In block 46, the audio streaming component receives two private communication channels from the remote site. These are routed in block 47 to the associated headphone playback via the audio interface output.

[0059] As background, an exemplary form of the audio streaming component of this system provides low-latency carrier of pulse code modulation (PCM) audio, such as SoundJack®. However, this system can also function as an insert between the input / output buffers of an audio streaming application and the system's capture / playback buffer, and can be configured for use with any suitable network music system that follows a low-latency streaming method.

[0060] The audio streaming method follows a general User Data Protocol (UDP) streaming design as described in XU, Aoxiang & Cooperstock, Jeremy (2002) "Real Time Streaming of Multi-channel Audio Data over Internet" 5120(I-3), or CHAFE, Chris, Scott Wilson, Randal J. Leistikow, Dave Chisholm, and Gary P. SCAVONE (2000) "A Simplified Approach To High Quality Music And Sound Over IP". In particular, the application used for system development, SoundJack®, benefits from the web GUI and server management of session metadata. An overview of the UDP audio streaming method is shown in Figure 8.

[0061] As background, the JACK® audio connection kit (JACK is a recursive acronym), which can provide real-time, low-latency connections for both audio and MIDI data between applications, was used for routing between applications within the system according to the present invention. This makes it possible to host audio processes in DAWs such as Reaper®. JACK v1.9.10 was used, and the connection between the hardware buffer and the application is<jack connect> and<jack disconnect> It was established or removed using the command. In particular, audio applications used with JACK must set the audio device to Jackrouter in the application settings.

[0062] As background, while Reaper™ was selected as the DAW to implement this invention, many alternative solutions are possible. DAWs can generally host functions such as channel routing, channel gain, channel summation, and convolution. The convolution process itself was performed using the measured HRTF and XMCFX™ convolver VST plugins. This convolution process uses Fast Fourier Transform (FFT) superimposed summation, with the first partition computed in the time domain to provide zero-latency throughput. For example, as shown in Figure 9, for each output frame y(n), the first partition should be computed in the time domain up to the superimposed region. The number of terms (samples) in the superimposed region can be calculated based on the length of the impulse response h(n). The superimposed region from each y(n) frame can be computed using the FFT method. Alternatively, all y(n) can be computed using the FFT method while allowing at least one audio buffer for process latency. It is worth noting that the implementation of this system can combine each of the functions provided by the third-party examples mentioned herein onto a single platform.

[0063] Variations of the present invention may include providing a predetermined virtual space in which the room ambience is measured / known, and the user / controller can select where the performer will be positioned, thereby assigning an appropriate headphone mix to that performer.

[0064] The position for measurement can be selected as the "best-sounding position" for application as a convolutional reverb, or it can be understood that the user can choose an entirely different location, such as a famous theater, to set up the performance. In a further alternative, a completely artificial reverb can be used to simulate the acoustic environment.

[0065] In certain forms, there may be provisions that allow each user to adjust their personal monitor mix to any desired position, including the positions of fellow musicians within the same local space.

[0066] This system and method can be summarized as a concerto-style musical performance tool that allows performers in a first spatial location and a second spatial location located away from the first spatial location to experience the perception that they are in the same location. The method / system requires acquiring at least one binaural room impulse response of a desired space (which may or may not be one of those locations), transmitting low-latency audio streams of performances in individual spatial locations over a network, and applying the binaural room impulse response (BRIR) as a real-time filter. In this way, sound sources from remote locations are perceived as being in the desired space when played through headphones. In one embodiment, one or both of the spatial locations can be divided into zones, and different BRIRs are applied depending on the zone location corresponding to the location where the BRIR was measured.

Claims

1. A method for concerto musical performance between a performer in a first spatial space having at least one zone in which performers are positioned and a performer in a second spatial space, A step of obtaining a virtual loudspeaker configuration for at least the first spatial space and / or another space in order to provide the acoustic characteristics of a desired space, wherein the virtual loudspeaker configuration is obtained by binaural acoustic measurement from at least two loudspeakers to the center of at least one zone, The steps of transmitting audio signals of one or more sound sources generated within the at least one zone to the second location space via a low-latency audio stream over a communication network, and the reverse step, The steps include: upon receiving the audio signal in the first and / or second location space, respectively, applying the virtual loudspeaker configuration as a real-time filter so that when played back through headphones, the sound source is perceived as being in the desired space; Methods that include...

2. The method according to claim 1, wherein the first position space is divided into a plurality of zones so that the obtained virtual loudspeaker configuration can be assigned to each zone based on the corresponding positions where the binaural acoustic measurements were performed.

3. The method according to claim 2, wherein each zone is associated with one or more performers.

4. The method according to any one of claims 1 to 3, wherein the step of obtaining the virtual loudspeaker configuration includes recording binaural acoustic measurements in at least the first spatial space, or creating / acquiring a corresponding recording in another space.

5. The method according to any one of claims 1 to 4, wherein the received audio signal is combined with one or more audio signals obtained from a local sound source.

6. The method according to claim 5, wherein a diffuse sound field measurement is obtained and applied as an impulse response to one or more audio signals obtained from a local sound source.

7. The method according to any one of claims 1 to 6, further comprising the steps of transmitting a video stream captured in the first location space to the second location space over a communication network, and vice versa, and displaying it on a display device.

8. The method according to claim 7, wherein the arrangement of the one or more sound sources in the stereo mix and applied together with the virtual loudspeaker configuration corresponds to the visual arrangement of the sound sources in the displayed video stream.

9. The method according to claim 7 or 8, wherein the display device is configured for split-screen visuals or has a plurality of adjacent display screens.

10. A system for collaborative musical performance between a performer in a first position space and a performer in at least a second position space, configured to perform the method according to any one of claims 1 to 9, wherein each position space is Voice acquisition device, At least one pair of headphones, Audio interface and A processor configured for audio processing and for sending and receiving low-latency audio streaming over a communication network, and Includes, The audio processing system includes applying a binaural room impulse response as a real-time filter to an input low-latency audio stream to implement a virtual loudspeaker configuration such that the captured sound source, when played back by the at least one pair of headphones, is perceived as being in a desired space.

11. The system according to claim 10, wherein each position further includes a camera and a display device, and the at least one processor is further configured for video streaming over a communication network.

12. The system according to claim 11, wherein the arrangement of the capture sound source in the stereo mix and applied together with the virtual loudspeaker configuration is configured to correspond to the visual arrangement of the sound source in the displayed video stream.

13. The system according to any one of claims 10 to 12, wherein there are a total of three or more location spaces, and the system is configured to stream captured sound sources between the location spaces as stereo streams for each location space.

14. The system according to claim 13, wherein an input stereo stream from a remote location space is summed with a locally captured sound source.

15. A method for obtaining a virtual loudspeaker configuration for a desired space, The steps include performing a first binaural acoustic measurement on a first loudspeaker at the center of at least one zone within the desired space, The steps include performing a second binaural acoustic measurement at the center of the at least one zone with respect to a second speaker positioned away from the first loudspeaker, The steps include preparing a binaural intra-room impulse response (BRIR) from the first and second measurements, and Methods that include...

16. The method according to claim 15, wherein the first loudspeaker is a left loudspeaker, the second loudspeaker is a right loudspeaker, and there is a third central loudspeaker from which a third binaural acoustic measurement is performed at the center of at least one zone in the desired space.

17. The method according to claim 15 or 16, wherein there are at least three zones within the desired space, and binaural acoustic measurements are repeated for each loudspeaker at the center of each zone.

18. The method according to any one of claims 15 to 17, wherein the desired space is a first spatial space where performers are positioned and / or another space such as a famous venue.