Novel time domain transfer function producing a virtual spatial sound field with restored ambience level

The method generates a virtual spatial sound field using calculated virtual sound sources to embed spatial information, addressing the lack of immersion in mono and stereo systems, achieving a realistic sound experience with ordinary speakers or headphones.

WO2026155670A1PCT designated stage Publication Date: 2026-07-23BÖHMER INNOVATIONS AB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BÖHMER INNOVATIONS AB
Filing Date
2025-12-11
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Current sound reproduction systems, particularly those using mono and stereo formats, lack embedded spatial information, leading to a lack of immersion and clarity in recreating real-life spatial sound fields, as they rely on room reflections to create a perceived soundstage, which is not present in the recording.

Method used

A method and apparatus that generates a virtual spatial sound field using at least two physical sound sources and additional virtual sound sources, calculated based on frequency response, to embed spatial information in the sound output, allowing for a 360x360 immersive experience without requiring multiple speakers or large rooms.

Benefits of technology

The virtual spatial sound field enhances immersion and clarity by embedding spatial information in the sound output, creating a realistic sound environment using ordinary stereo speakers or headphones, without the need for complex setups or ideal acoustic rooms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2025010065_23072026_PF_FP_ABST
    Figure SE2025010065_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides a method comprising the following steps: - providing at least two physical sound sources, preferably at least two stereo loudspeakers or headphone transducers, providing a physical sound output; - providing at least two virtual sound sources based on a calculation of frequency response, said at least two virtual sound sources each providing delayed virtual sound output based on the physical sound output from said at least two physical sound sources; and - producing a virtual spatial sound field with restored ambience level based on the physical sound output from said at least two physical sound sources, by using a digital sound processing unit.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Novel time domain transfer function producing a virtual spatial sound field with restored ambience level.

[0002] Background

[0003] From the very beginning it has been the goal of scientists that started recording sound and subsequently reproduce it to create a life like copy of the original sound event. Over the years, sound quality and immersion have gradually improved. At the start using only one channel, mono, recording and reproduction practices in the fifties moved to two channels, stereo. In the seventies, Dolby introduced Dolby Stereo, which despite its name was the first successful multi-channel surround sound format. In more recent years, object-based audio formats have emerged, increasing the level of immersion and number of channels even further. Surround sound is now prevalent in movies and video games. Although multi-channel sound distribution and reproduction have been available to consumers for a long time, stereo, however, remains the dominant sound reproduction format. A commercial cinema built to reproduce one of the object-based multi-channel formats can sound impressive and approach the spatial sound field of a real-life event, but the required number of speakers located around the room can run into the hundreds. Such a setup is far beyond what any normal consumer would think tolerable in a home environment both for reasons of cost and aesthetics. For consumers, the number of accepted playback channels appears to have essentially maxed out at two.

[0004] Renderers that aim to reproduce multi-channel sound on reproduction equipment utilizing two channels have thus far not been successful in doing so. Single box Bluetooth speakers, soundbars, stereo speakers and headphones are unable to reproduce sound that is approaching a real-life spatial sound field. A commercial cinema sound system or a high-end consumer home theatre setup using perhaps nine or more speakers, being significantly superior in recreating a real-life immersive experience. The reason a multitude of speakers are required is that Mono, stereo or multi-channel recordings do not have spatial information embedded in the signals fed to the speakers, the direction of the sound in 360x360 space, and thus need speakers in many locations and a room to create a spatial sound field that mimics a real-life sound field. The lack of embedded spatial information also causes another problem; for the reproduction to sound acceptable a significant reduction ofreverberant or indirect sound becomes necessary even further removing the reproduced sound from a real-life experience.

[0005] Throughout our lives we are surrounded by sound waves travelling through the air around us. Even in quiet places there is a plethora of sounds forming dense and complex sound fields. Sound waves are continuously and simultaneously arriving at each moment in time from diverse sources and directions to our ears forming a continuous soundtrack that never stops. Sound waves behave much like water waves. Think about a pool with ten people, each person a sound source, splashing around making small waves. The waves propagate across the pool’s water surface from different directions and arrive at our location in the pool at different points in time, some having bounced on pool walls and other objects. Water waves do however only propagate across the water surface, a two-dimensional space, whereas waves in a sound field are traveling in three dimensions. Our ears and brain are left with a formidable task to interpret and make sense of even the simplest of sound fields occurring in our environment.

[0006] Sound fields consist of direct sound waves arriving directly from the sound source to our ears and indirect sound waves that have at least bounced on one physical object before reaching us. Most sound waves are indirect, having bounced on objects many times before arriving i.e., more sound energy is present in indirect waves compared to direct waves. Sound waves have three fundamental properties that humans require to be able to interpret the sound field around us: arrival time, direction, and momentary amplitude. The human hearing “apparatus” consists of our two ears sensing the arrival times and momentary sound pressure, together with the physical properties of our ears, head and torso providing the means to decode directionality, all input at the end interpreted by our brain.

[0007] The brain tries to make sense of sound fields in much the same way as our vision work. It simplifies the sound field by grouping direct and indirect sound waves from a single source into sound objects. As an example, the sound we experience and call a doorbell is the direct sound of the doorbell, together with all its attendant indirect sound in situ. We do not experience direct sound and indirect sound separately; our brain automatically and unconsciously groups the doorbell sound waves togetherinto an object. For the sound grouping process to work, which is easy to imagine is very complex, information about all the properties of the sound waves in the sound field is required: arrival time, sound wave direction and momentary amplitude.

[0008] Current recording techniques capture the arrival time and momentary amplitude, the sound pressure, of the sound field in one or more microphone locations. Directional spatial information, i.e., the three-dimensional direction of the sound waves in those locations is not captured. Although microphones exhibit some directionality and captures sound differently depending on direction, this directionality cannot be translated into any meaningful information by the human brain.

[0009] Single channel mono playback of sound can provide a soundstage with some perceived depth and height. The presence of a soundstage implies that spatial information is available, but mono has no spatial information embedded in the recorded sound and the limited soundstage that is presented is created by reflections from surfaces in the listening environment. The loudspeaker and room together render a spatial sound field, but the spatial information is not in the recording. The spatial sound field creates the perception of a cloud of sound around the single loudspeaker source. The absence of spatial information in the recording is easily verified by listening to mono in an anechoic environment. In the anechoic environment the soundstage cloud disappears.

[0010] Two channel stereo provides some spatial information and can produce a relatively continuous horizontal plane of sound between the two loudspeakers with a bit of soundstage height and depth. The spatial information embedded in a stereo recording is however limited to left and right localization of sound sources on a single plane between the loudspeakers. The localization is created by arrival time and amplitude differences between the two speakers. Perceived soundstage height and depth are again created by reflected sound in the listening environment. The loudspeakers together with reflecting surfaces within the environment render a spatial sound field, the soundstage, located between the speakers possessing limited height and depth. The sound field present in the room is, however, different from the spatial sound field at the recorded event, it is rendered independently by the speakers and room. Without these room reflections the soundstage disappears, which can be demonstrated using a pair of highly directional speakers, i.e. parabolic,or speakers in an anechoic room. There is no spatial information embedded in the recording and the soundstage disappears when reflections from the environment are removed. This is further evidenced by listening to stereo recordings using headphones, the sound is invariably located inside the listener’s head.

[0011] The lack of spatial information in recordings hinders grouping of direct and indirect sound from sources into sound objects. This leaves unattached indirect sound that does not belong to a certain direct sound. There is significantly more energy in the indirect sound and the natural attenuation of indirect sound belonging to a certain direct sound that occurs in the brain, to automatically put focus on the direct sound and its location, does not work anymore. This leads to recordings that sound excessively reverberant and very strange if containing the same amount of indirect sound as is present in real life. Therefore, to balance the perceived amount of direct and indirect sound in recordings, recording engineers routinely place microphones close to performers on stage, not out in the hall where the audience normally is located. The microphones consequently capture less indirect sound and more direct sound due to the proximity to the source.

[0012] Investigations of the ratio between indirect reflected sound and direct sound reveals that there is typically 12dB less indirect sound in recordings compared to representative sound experienced in a concert hall. The significant lack of indirect and thereby reverberant information is clearly audible and inherent to the recording process, thus being present throughout all genres of recorded music, not just live music in concert halls. Due to the human inability to group direct and indirect sound in recordings into objects, attenuation of reverberant sound is required to prevent recordings from sounding excessively reverberant, but the resultant sound has a significant lack of reverberant energy and does not sound like the real thing.

[0013] We know from psycho acoustic research that human hearing is extremely sensitive to time domain properties of sound and uses time domain information extensively to figure out the contents of sound fields. Location of sound sources is determined in milliseconds when the first direct sound wave arrives, size and proportions of the sound source are assessed during the following roughly ten milliseconds by the spatial properties of the early indirect sound waves. After about 25ms, the indirectsound waves are beginning to be heard increasingly as echoes, the reverberation tail of the sound, lasting much less than a second in smaller rooms up to many seconds in a large church.

[0014] The missing 12dB indirect sound in recordings could in theory be added and rendered by the speakers and room in playback but unfortunately this does not work well in domestic environments. In homes, room sizes are normally below 100m2 or at least not significantly larger. The sound fields produced within domestic room sizes will contain first indirect reflections from the room arriving in the range 0-15ms after the direct sound from the speakers. In smaller rooms the first reflected indirect sound arrives earlier than it does in a larger room. From concert hall acoustic research, we know that reflections that arrive before approximately 25ms will make the sound of the instruments less clear and add a harsh flavor to the sound. Sound that arrives after 25ms and later will increasingly add beauty to the sound. Later arriving sound reflections are desirable but the early reflections from the stage area are unfavorable. Similarly, the first indirect reflections from a domestic room arriving in the range 0-15ms does not add beauty to the sound, it produces a harsh and unclear sound. In addition, most loudspeakers are relatively directional and do not propagate enough energy into the room in other directions than forward to be able to add all the missing 12dB of indirect sound.

[0015] An omni directional speaker and a room can produce a sound field with enough indirect energy, but the sound will be very unclear and harsh due to early arriving indirect reflections. Any non-ideal acoustic properties of the room will also heavily color the sound. A pair of superb omni directional speakers in a room with excellent acoustic properties and a size of at least 250m2 with significant ceiling height can render a passable spatial sound field with an appropriate amount of indirect sound energy. In such a large room the earliest indirect reflections arrive later than 25ms, which adds to the experience of the sound and makes it possible to fill in the missing 12dB indirect energy. However, the room will have a significant reverberation of its own that will be overlayed on everything that is played back. Some recordings may benefit from the room’s particular acoustic properties but quite a few may not do so. A further obvious issue is that few have the very significant budget and space needed to acquire a 250m2 size room with high ceiling specifically designed topossess excellent acoustics. A more accessible and practical solution that works on all music without overlaying its own reverberant fingerprint and functioning in normal environments is clearly desirable.

[0016] The presented invention reveals a solution, it is a method and apparatus that produces a virtual spatial sound field. The virtual spatial sound field contains sound emanating from many locations in virtual 360x360 space around the listener, it has the indirect sound organized so that it arrives at appropriate points in time not causing harshness or blurring. Indirect reflections are properly diffused to prevent hard reflections, the virtual reflections have “bounced” on acoustically ideal surfaces, the reflected sound is decorrelated to avoid room effects like boomy lower frequencies and does not require any additional virtual room related reflections or reverberation, thus avoiding any sameness to the reverberation tail of sounds. Since the virtual sound field contains spatial information required for grouping to work, it can also contain the required amount of indirect sound to replicate a real-life sound field. The virtual spatial sound field can be used with ordinary stereo recordings. It can also be used with renderers producing sound for stereo playback. The renderers’ outputs are complementary to and compatible with the virtual spatial sound field. The renderers can be multi-channel down mixers, game play renderers, virtual reality renderers or similar for both speakers and headphones.

[0017] The reproduction of the virtual sound field does not require a complex setup with many speakers, nor does it require very large, acoustically ideal rooms. A virtual spatial sound field can be created by merely employing two forward radiating ordinary stereo speakers in a normal domestic room. The virtual spatial sound field embeds psycho acoustic time and frequency information in the signals fed to the two usual left and right speakers that makes the human brain believe in the existence of an immersive 360x360 degrees real-life spatial sound field, the spatial information is however only virtually present.

[0018] The virtual spatial sound field works well with ordinary forward radiating loudspeakers, but the virtual spatial sound field is mixed with and influenced by the sound field rendered by the speakers and playback room. A reverberant room combined with wide dispersion speakers will introduce more sound from the roomwith, as discussed earlier, its less ideal sound field properties and the result will not be quite as good as utilizing speakers with controlled directionality in a less reverberant room.

[0019] Two loudspeakers in a room playing back ordinary stereo signals will need to spread sound in the room to produce a spatial sound field. When more sound is distributed around the room by the speakers, the sound stage is perceived to be larger but at the same time if too much sound is distributed the risk of the sound becoming unclear and harsh increases, so a balance needs to be struck. Utilizing the virtual spatial sound field, the spatial sound stage is already present in the feed to the speakers, so it is no longer necessary to use the room to render the spatial components of the sound field. Thus, a reduction of sound distributed around the room by the speakers is desirable and speakers with controlled dispersion, wide enough to cover a listening area but not wider, produces a better result. The room interactions are thereby reduced and the colorations that occur due to the acoustically less ideal properties of the room can be decreased without influencing the perceived sound stage negatively. The virtual spatial sound field is immersing the listener in what is heard as a 360x360 sound field that is both fully enveloping and significantly larger than a stereo soundstage traditionally produced by speakers and a room.

[0020] In relation to the above it should be mentioned that the method according to the present does not only refer to the stereo alternative with two channels for the insignals. Also multiple discrete channels having signal processing already incorporated may also function according to the present invention.

[0021] The virtual spatial sound field can also be reproduced over headphones, but the perception of the sound field is different from loudspeakers. The word headphone is used to mean any type of possible headphone device, over ear, on ear, in ear, ear buds etc. Over headphones, the sound will be perceived to be located 360x360 degrees around the head of the listener, just outside of the head. The reason the sound field is perceived to be limited to just outside of the listener’s head is the missing time domain acoustic information created by the human body when it is located within a sound field. With headphones, the body is obviously not in a soundfield, the sound is fed directly into each ear without being influenced by the body. Due to the missing time domain acoustic information normally created by the body, the sound field ends up being perceived to be around the listener but limited to just outside the head and not farther away like it would be with loudspeakers. To create a virtual sound field with headphones that sounds like the virtual sound field produced with loudspeakers the presented invention can be combined with the solution revealed by PCT / SE2021 / 051005, 2021-10-14, “Sound reproduction with multiple order HRTF between left and right ears”.

[0022] Detailed description of the invention

[0023] According to a first general aspect of the present invention, there is provided a method comprising the following steps:

[0024] - providing at least two physical sound sources, preferably at least two stereo loudspeakers or headphone transducers, providing a physical sound output;

[0025] - providing at least two virtual sound sources based on a calculation of frequency response, said at least two virtual sound sources each providing delayed virtual sound output based on the physical sound output from said at least two physical sound sources; and

[0026] - producing a virtual spatial sound field with restored ambience level based on the physical sound output from said at least two physical sound sources,

[0027] by using a digital sound processing unit.

[0028] Based on the above, according to one embodiment of the present invention, there is provided a method comprising the following steps:

[0029] - providing at least two physical sound sources, preferably at least two stereo loudspeakers or headphone transducers, providing a physical sound output;

[0030] - providing at least two virtual sound sources based on a calculation of frequency response, said at least two virtual sound sources each providing delayed virtual sound output based on the physical sound output from said at least two physical sound sources; and

[0031] - producing a virtual spatial sound field with restored ambience level in relation to the physical sound output from said at least two physical sound sources,

[0032] by using a digital sound processing unit,wherein the step of providing at least two virtual sound sources involves the following steps for each virtual sound source:

[0033] - calculating a minimum phase frequency response deviation from Head Related Impulse Response data (HRIR) data for a real speaker’s location, the left and right speakers, respectively, preferably the front left and right speakers, respectively; - calculating a minimum phase frequency response deviation using HRIR data for a location of the virtual sound source; and

[0034] - subtracting a calculated frequency response for a location of the real speaker from the frequency response of the location of the virtual sound source, preferably subtracting a left and right real speaker’s minimum phase frequency response deviation from a virtual left and right surround speaker’s minimum phase frequency response deviation, respectively, to obtain the complete aggregate frequency response deviation for the virtual left and right surround speakers.

[0035] Furthermore, according to yet another embodiment of the present invention, there is provided method comprising the following steps:

[0036] - providing at least two physical sound sources, preferably at least two stereo loudspeakers or headphone transducers, providing a physical sound output;

[0037] - providing at least two virtual sound sources based on a calculation of frequency response, said at least two virtual sound sources each providing delayed virtual sound output based on the physical sound output from said at least two physical sound sources; and

[0038] - producing a virtual spatial sound field with restored ambience level in relation to the physical sound output from said at least two physical sound sources,

[0039] by using a digital sound processing unit,

[0040] wherein the method involves using headphones where the headphone transducers are physical sound sources and wherein the step of providing at least two virtual sound sources involves the following steps for each virtual sound source:

[0041] - calculating a minimum phase frequency response deviation from Head Related Impulse Response data (HRIR) data for locations of the headphone transducers, left and right, respectively;

[0042] - calculating a minimum phase frequency response deviation using HRIR data for a location of the virtual sound source; and- subtracting a calculated frequency response for a location of the headphone transducer from the frequency response of the location of the virtual sound source, for obtaining the complete aggregate frequency response deviation for the virtual sound source.

[0043] In relation to the expressions “ambience level” and “restored ambience level”, it may be mentioned that these expressions are standard to use in this technical field.

[0044] When recording engineers use the term "Ambience level" it is normally used to refer to the amount of sound from the environment in a recording venue separated from direct sounds from the sound sources on stage. "Ambience" encompasses reverberant indirect sound and any additional sounds in the venue such as clapping sounds from and audience etc., and the term "Ambience level" denotes the sound intensity of these indirect or background sounds.

[0045] It should be noted that the method according to the present invention may be a computer-implemented method or implemented in similar software-oriented devices, such as any form of digital sound processing units, both for online and offline processing, and combinations thereof, with or without the use of the cloud etc.

[0046] An apparatus with a microcontroller or digital signal processor of some sort that is capable of processing sound is suitably used. The apparatus can process sound online as it is being streamed in real time, for example while being recorded, broadcast, or played back. It can also process recorded sound off-line, for example computer files containing recoded audio that are processed and written back to disk, or whatever storage medium that happens to be in use. The processing apparatus can also be a cloud server, processing audio over the internet. As examples, it can be a service processing off-line contents sent to it, or it can be an on-line service processing audio data as it is streamed.

[0047] Ignoring some differences that will be explained later, the core data processing method is the same regardless of whether the target playback devices are speakers or headphones. The method works by fashioning virtual sound sources located in several places in 360x360 space around the listener that operates in combination with the main physical left and right channels. The individual sound sources and locations are being used together to generate the virtual spatial sound field. Themethod does not require any additional virtual room related reflections or reverberation, thus avoiding sameness to the reverberation tail of sounds.

[0048] As a minimum 4 sound sources are needed, two physical sources, left and right, together with two virtual surround sources left and right back located approximately in the traditional surround locations. Preferably at least 9 sound sources should be used, 7 virtual and 2 real sources, spread out like a virtual surround sound system with virtual speakers surrounding the listener. An extended system can preferably comprise 18 virtual sound sources in addition to the 2 real speakers. There is no maximum number of virtual sound sources, it could be as many as 100, but audible gain diminishes gradually above 20. As many virtual sound sources as 100 increases the required sound processing capacity whilst just providing small gains above 20. When considering real and virtual sound sources conceptually as separate channels it is important to keep in mind that although the virtual sound sources are separated from the real main left and right speaker outputs in the theoretical discussion, all the sound is played back through the real speakers, no other real sound sources are required.

[0049] In front of the listener there are the three direct sound sources, left, center and right (LCR). The center channel is a virtual channel and not mandatory but makes it possible to produce a better tonal balance for sounds located in the middle of the stage and is therefore desirable. The other virtual sound sources are used to render the spatial sound field and the sound from these sources is delayed in relation to the direct sound from the LCR channels. The delayed virtual channels preferably include all or some of Front LR floor, Front LR ceiling, Side Surround LR, Side Surround LR floor, Side Surround LR ceiling, Rear Surround LR, Rear Surround LR floor, Rear Surround LR ceiling and center back; in total 20 sound sources, 2 real and 18 virtual. The delay in the sound output through the virtual sound sources used to render the spatial sound field is long enough to prevent loss of clarity and avoid harshness but not so extended that the spatial field produces an added reverberation effect.

[0050] Preferably the delay is in the range 15ms to 75ms, more preferably 20ms to 50ms, most preferably 25ms to 45ms. The sound from the delayed virtual sound sources should also be decorrelated in relation to the LCR channels to avoid undesirable “room” resonance effects due to correlated summation between the different soundsources. The sound from the delayed sources should also be diffused in time to virtually emulate acoustic diffusion. The virtual acoustic diffusers should preferably use stochastic diffusion and / or diffusion based on a number-series derived using the golden ratio. Decorrelation can be implemented using all-pass filters and diffusion by employing FIR filters. There are obviously other means available that will produce adequate diffusion and decorrelation, the ones mentioned are only examples. The emulated acoustic diffusers should release the diffused energy from 0 to 0.1ms up to 3ms, more preferably from 0 to 0.15ms up to 0.5ms. The sound output level through each delayed virtual sound source is preferably attenuated by -3dB to -30dB, more preferably in the range -6dB to -20dB.

[0051] The delayed virtual sound source outputs also produce low level early reflections at levels in the level range -40dB to -90dB relative to the main LCR channels, preferably in the level range -60dB to -80dB. The early reflections are necessary to avoid unnaturalness to the sound and are there only to complement the main delayed energy. The delay of the early reflections is in the range of 5ms to 20ms related to the LCR outputs. The low level of these early reflections does not cause unclarity nor any harshness.

[0052] The virtual center channel’s direct sound output is produced by calculating the L+R signal. The delayed virtual sound source outputs are derived from the left and right input signals by performing left minus right (L-R) and right minus left (R-L) calculations. The L-R signal contains the space to the left of the left speaker and the R-L signal covers the space to the right of the right speaker. The two derived signals predominantly contain indirect ambient sounds and reverberation. The signals are preferably amplified so that all virtual sound source outputs summed together produce amplification of the indirect ambient sounds and reverberation by 6dB to 18dB, more preferably by 9dB to 14dB and most preferably by 12dB in relation to their original derived level. The L-R signal is fed to the left side of the delayed virtual sound source outputs and the R-L to the right side. The delayed virtual sound source outputs are in addition to these main delayed signals, L-R in the left channel and R-L in the right channel, producing the formerly mentioned low-level early reflection part. The early reflection part is calculated in the same way, L-R and R-L, but the R-Lsignal is supplied through the left side of the delayed virtual sound source outputs and the L-R signal though the right side.

[0053] A sound wave, direct or indirect, arriving to the human body from any point in 360x360 space will initially exhibit a minimum phase frequency deviation and then become spread out in the time domain by delayed sound bouncing from the body, causing further non minimum phase frequency response deviations. Each source point in space has a minimum phase frequency response deviation associated with it. This minimum phase frequency response deviation can be calculated from standard Head Related Impulse Response data (HRIR) by removal of the nonminimum phase delayed parts of the HRIR then calculating the frequency response from the minimum phase part of the HRIR. Note that this is different from calculating the frequency response from the entire HRIR response and the calculated minimum phase frequency response is different compared to the ordinarily used frequency deviation calculated from the entire HRIR data that includes the non-minimum phase parts as well. All minimum phase frequency deviations should be obtained by averaging minimum phase frequency deviations calculated using HRIR data from several individuals. Averaging across many individuals is important for the results to be generally applicable and work well for a large majority of people. Averaging between as few as 5 individuals will improve the result significantly but much better results can be obtained if at least 50 to more than 200 are used.

[0054] To create a virtual sound source the complete aggregate frequency response associated with the virtual sound source location in 360x360 space must be calculated. The first step is to calculate the minimum phase frequency response deviation from HRIR data for the real speaker’s location, the front left and right speakers. Only the first arrival minimum phase frequency response deviation should be used, later arriving excess phase deviations caused by delayed sound reflections on the human body should not be included as those will inevitably be added by the body of the listener and should not be duplicated. The second step is to calculate the minimum phase frequency response deviation using HRIR data for the virtual sound source location. The third step is to subtract the calculated frequency response for the real speaker location from the frequency response of the virtual speaker location, as an example the left real speaker’s minimum phase frequency response deviationis subtracted from the virtual left surround speaker’s minimum phase frequency response deviation to obtain the complete aggregate frequency response deviation for the virtual left surround speaker.

[0055] For headphones, the minimum phase frequency response deviation from HRIR data for the real front speakers should not be used as described above since the real speakers are now the left and right transducers in the headphones. The first step is thus changed and the calculated minimum phase frequency response deviation for the real speaker’s location is the location of the left and right transducers in the headphones. The frequency response for the headphone transducers are now the real speakers, the front left and right speakers become virtual speakers and should consequently have an associated virtual complete aggregate minimum phase frequency response deviation added, like all other virtual sound sources.

[0056] Similarly, the virtual spatial sound field of course works with real loudspeakers located in other positions than the traditional front left and right locations mentioned previously above and the calculations for use with loudspeakers in alternative locations should be modified in the same manner as just described for headphones.

[0057] A further final processing step is also required for headphones, which should not be used with loudspeakers. When the human body is in a sound field, sound waves bouncing on and circulating around various body parts will cause delayed sound to arrive at the ears, delayed compared to the direct sound arriving without delay into the ears. When the real sound sources are loudspeakers and the listener’s body is in the sound field from the speakers this delayed part of the transfer function is already present due to the human body being in the sound field and thereby inescapably generating these delayed parts of the response. Over headphones this part must be virtually added.

[0058] Picture 1 illustrates sound paths around the human body from a sound source in space to and around a listener’s head. Number 1 is the listener, 2 the sound source and 3 to 8 are visualized sound wave paths to and around the head. The picture only illustrates one sound source location but any location in three-dimensional space has a similar set of imaginable sound paths associated with it. These sound paths must be added to a virtual location over headphones for the perception of properpositioning of the sound, outside the body in the same location as they would have been heard using speakers. Without this delayed part of the transfer function, emulating the human body being present within a sound field, the sound is perceived to be located just outside the head.

[0059] Each sound path 3 to 8 has a time delay, a frequency response and attenuation associated with it. Path 3 has a time delay, the travel time of sound from the sound source 2 to the right ear but in this special case, since this is the first arrival of sound to the listener, the delay is zero as there is no need to have a delay that parallels the sound travel time to reach the listener. Attenuation in this specific first order path is also zero since the sound travels directly to the ear without any obstacles that can produce attenuation. The sound wave will however not stop when it has reached the right ear. It will continue along path six around the head to the left ear. This path has an interaural time delay due to sound travel time, a frequency response due to the shadowing of higher frequencies by the head and attenuation caused by the travel around the head to the other ear. When the sound wave has reached the left ear, it will again continue to travel along path eight back to the right ear and once more this path has a time delay, a frequency response and attenuation associated with it. For reasons of clarity the picture does not illustrate higher order paths, but the principle should now be obvious, and it is easy to extrapolate any higher order paths by just continuing with them around the head.

[0060] The sound source 2 in Picture 1 can be both a direct sound from a sound source or indirect sound, having bounced on one or more objects. Each sound wave arriving at the human body, irrespective of whether it is direct or reflected sound, will generate the illustrated type of sound paths around the human body.

[0061] Embodiments of the invention

[0062] Below some embodiments of the present invention are provided and discussed further.

[0063] According to one embodiment, all sound is played back through said at least two physical sound sources. Moreover, in relation to the case with stereo loudspeakers, according to one embodiment, said at least two physical sound sources represent aleft channel and a right channel, respectively. Furthermore, according to yet another embodiment, said at least two virtual sound sources represent one or more left back and one or more right back virtual sound sources, respectively, preferably said method involves providing multiple virtual sound sources located at different places in a 360x360 space around a listener, in which multiple virtual sound sources operate in combination with said at least two physical sound sources representing at least a left channel and a right channel.

[0064] As mentioned above, the present invention also embodies alternatives to the two stereo input channels. According to one embodiment, said input are multiple discrete channels having signal processing already incorporated. In line with both these options, according to one embodiment input signals are provided via two stereo input channels or as multiple discrete channels having signal processing already incorporated.

[0065] As explained above, the method according to the present invention is time-oriented, meaning that a delay in the sound output is used. In line with this, according to one embodiment of the present invention, a delay in a sound output through said at least two virtual sound sources, based on a time domain response, and used to render the spatial sound field is in a range of from 15ms to 75ms, preferably in a range of from 20ms to 50ms, more preferably in a range of from 25ms to 45ms.

[0066] Moreover, the delayed sound output is suitably handled by different means according to the present invention, First of all, according to one embodiment of the present invention, the delayed virtual sound output is decorrelated in relation to the physical direct sound sources, such as LR channels, to avoid undesirable “room” resonance effects due to correlated summation between the different sound sources, e.g. by using all-pass filters. In relation to the above it may be said that “direct sound” refers to the portion of a sound wave that travels direct from the source to a listener without reflecting off any surfaces. Moreover, it may also be mentioned that “room” resonance can be regarded as the acoustic phenomenon where sound waves at specific frequencies are naturally amplified or attenuated due to the physical dimensions and shape of an enclosed space.Secondly, according to yet another embodiment, the delayed virtual sound output is diffused in time to virtually emulate acoustic diffusion, e.g. by employing FIR filters, preferably the delayed virtual sound output is both decorrelated and diffused.

[0067] Decorrelation and diffusion are not one and the same. Acoustic decorrelation refers to the process of reducing the similarity (or correlation) between two or more sound signals in terms of their phase and timing. Acoustic diffusion on the other hand refers to the process of scattering sound waves in different directions to reduce acoustic anomalies like echoes, standing waves, and flutter echoes. This process helps create a more even distribution of sound energy within a space, leading to improved sound quality and clarity. According to the present invention, decorrelation is thus used to prevent virtual signals to be in phase with real signals, and diffusion is used for distributing the energy over time, such as when subsequent reflection of sound is obtained. Diffusion and decorrelation may be used both individually and together. Moreover, some cases of diffusion may provide decorrelation as such, but this is not guaranteed. Furthermore, it may be noted that the particular frequency in question may influence when decorrelation and diffusion, respectively, are suitable. Normally, decorrelation is often suitable below 500 Hz and diffusion above 500 Hz.

[0068] According to yet another embodiment of the present invention, the delayed virtual sound output also produces low level early reflections at levels in a level range of -40dB to -90dB relative to the physical sound sources, such as LR channels, preferably in a level range of -60dB to -80dB.

[0069] Moreover, according to one embodiment, said step of providing at least two virtual sound sources comprises calculating a complete aggregate frequency response associated with a set of virtual sound source locations in a 360x360 space.

[0070] Furthermore, according to yet another embodiment, the method involves using as a minimum 4 sound sources of which two are physical sources, left and right, together with two virtual surround sources left and right back located approximately in the traditional surround locations, suitably behind a listener, preferably the method involves using at least 9 sound sources of which at least 7 are virtual sound sources spread out like a virtual surround sound system with virtual speakers surrounding the listener. In relation to the expression “traditional surround locations” it may bementioned that such traditional surround locations are front left right and center in front of the main listening position plus rear left right behind the main listening position, in total encompassing five speakers.

[0071] Moreover, and as mentioned above, according to one embodiment, the step of providing at least two virtual sound sources involves the following steps for each virtual sound source:

[0072] - calculating a minimum phase frequency response deviation from Head Related Impulse Response data (HRIR) data for a real speaker’s location, the left and right speakers, respectively, preferably the front left and right speakers, respectively; - calculating a minimum phase frequency response deviation using HRIR data for a location of the virtual sound source; and

[0073] - subtracting a calculated frequency response for a location of the real speaker from the frequency response of the location of the virtual sound source, preferably subtracting a left and right real speaker’s minimum phase frequency response deviation from a virtual left and right surround speaker’s minimum phase frequency response deviation, respectively, to obtain the complete aggregate frequency response deviation for the virtual left and right surround speakers.

[0074] In relation to the above, front left and right speakers are suitably used, but the left and right speakers may be located elsewhere than in front of the listener.

[0075] For the headphones alternative, according to one embodiment, the method involves using headphones where the headphone transducers are physical sound sources and wherein the step of providing at least two virtual sound sources involves the following steps for each virtual sound source:

[0076] - calculating a minimum phase frequency response deviation from Head Related Impulse Response data (HRIR) data for locations of the headphone transducers, left and right, respectively;

[0077] - calculating a minimum phase frequency response deviation using HRIR data for a location of the virtual sound source; and

[0078] - subtracting a calculated frequency response for a location of the headphone transducer from the frequency response of the location of the virtual sound source,to obtain the complete aggregate frequency response deviation for the virtual sound source.

[0079] In the case of headphones, the standard L and R locations have become the virtual sound sources. According to yet another embodiment, said method also comprises postprocessing including adding a delayed part of a transfer function as a virtual addition to the output sound in the headphones.

[0080] Furthermore, according to one embodiment, the method comprises using at least three direct sound sources, which are left, center and right (LCR), direct sound sources being such where the sound arrives directly from the sound source to someone’s ears, wherein the left and right are physical sound sources and the center channel is a virtual center channel providing a better tonal balance. Moreover, according to one embodiment, the other virtual sound sources are used to render the spatial sound field and the sound from these other virtual sound sources is delayed in relation to the direct sound from the LCR channels. Furthermore, according to yet another embodiment, a direct sound output of said virtual center channel is produced by calculating the sum of the signals to the left and right physical sound sources, preferably by performing left plus right (L+R) calculations on said signals. Moreover, according to yet another embodiment, difference signals, left minus right (L-R) and right minus left (R-L) of the left and right physical sound sources are amplified, preferably said signals are amplified so that all virtual sound source outputs summed together produce amplification of the difference signals by 6dB to 18dB, more preferably by 9dB to 14dB, and most preferably by 12dB in relation to their original derived level.

[0081] As hinted above, the method according to the present invention may be implemented in many different ways. According to one embodiment, the digital sound processing unit is arranged to process sound on-line as it is being streamed in real time, for example while being recorded, broadcast, or played back, or process recorded sound off-line, for example computer files containing recoded audio that are processed and written back to disk, wherein said digital sound processing unit is a computer implemented unit or other type of software unit, a cloud server or other processing audio over the internet, ora combination thereof.

Claims

Claims1. A method comprising the following steps:- providing at least two physical sound sources, preferably at least two stereo loudspeakers or headphone transducers, providing a physical sound output;- providing at least two virtual sound sources based on a calculation of frequency response, said at least two virtual sound sources each providing delayed virtual sound output based on the physical sound output from said at least two physical sound sources; and- producing a virtual spatial sound field with restored ambience level in relation to the physical sound output from said at least two physical sound sources,by using a digital sound processing unit,wherein the step of providing at least two virtual sound sources involves the following steps for each virtual sound source:- calculating a minimum phase frequency response deviation from Head Related Impulse Response data (HRIR) data for a real speaker’s location, the left and right speakers, respectively, preferably the front left and right speakers, respectively; - calculating a minimum phase frequency response deviation using HRIR data for a location of the virtual sound source; and- subtracting a calculated frequency response for a location of the real speaker from the frequency response of the location of the virtual sound source, preferably subtracting a left and right real speaker’s minimum phase frequency response deviation from a virtual left and right surround speaker’s minimum phase frequency response deviation, respectively, to obtain the complete aggregate frequency response deviation for the virtual left and right surround speakers.

2. A method comprising the following steps:- providing at least two physical sound sources, preferably at least two stereo loudspeakers or headphone transducers, providing a physical sound output;- providing at least two virtual sound sources based on a calculation of frequency response, said at least two virtual sound sources each providing delayed virtual sound output based on the physical sound output from said at least two physical sound sources; and- producing a virtual spatial sound field with restored ambience level in relation to the physical sound output from said at least two physical sound sources,by using a digital sound processing unit,wherein the method involves using headphones where the headphone transducers are physical sound sources and wherein the step of providing at least two virtual sound sources involves the following steps for each virtual sound source:- calculating a minimum phase frequency response deviation from Head Related Impulse Response data (HRIR) data for locations of the headphone transducers, left and right, respectively;- calculating a minimum phase frequency response deviation using HRIR data for a location of the virtual sound source; and- subtracting a calculated frequency response for a location of the headphone transducer from the frequency response of the location of the virtual sound source, for obtaining the complete aggregate frequency response deviation for the virtual sound source.

3. The method according to claim 1 or 2, wherein all sound is played back through said at least two physical sound sources.

4. The method according to any of claims 1-3, wherein said at least two physical sound sources represent a left channel and a right channel, respectively.

5. The method according to any of claims 1-4, wherein said at least two virtual sound sources represent one or more left back and one or more right back virtual sound sources, respectively, preferably said method involves providing multiple virtual sound sources located at different places in a 360x360 space around a listener, which multiple virtual sound sources operate in combination with said at least two physical sound sources representing at least a left channel and a right channel.

6. The method according to any of claims 1-5, wherein a delay in a sound output through said at least two virtual sound sources, based on a time domain response, and used to render the spatial sound field is in a range of from 15ms to 75ms, preferably in a range of from 20ms to 50ms, more preferably in a range of from 25ms to 45ms.

7. The method according to any of claims 1-6, wherein the delayed virtual sound output is decorrelated in relation to the physical direct sound sources, such as LR channels, where the sound arrives directly from the sound source to someone’s ears, to avoid undesirable “room” resonance effects due to correlated summation between the different sound sources, e.g. by using all-pass filters.

8. The method according to any of claims 1-7, wherein the delayed virtual sound output is diffused in time to virtually emulate acoustic diffusion, e.g. by employing FIR filters, preferably the delayed virtual sound output is both decorrelated and diffused.

9. The method according to any of claims 1-8, wherein the delayed virtual sound output also produces low level early reflections at levels in a level range of from -40dB to -90dB relative to the physical sound sources, such as LR channels, preferably in a level range of from -60dB to -80dB.

10. The method according to any of claims 1-9, wherein said step of providing at least two virtual sound sources comprises calculating a complete aggregate frequency response associated with a set of virtual sound source locations in a 360x360 space.

11. The method according to any of claims 1-10, wherein the method involves using as a minimum 4 sound sources of which two are physical sources, left and right, together with two virtual surround sources left and right back located approximately in the traditional surround locations, suitably behind a listener, preferably the method involves using at least 9 sound sources of which at least 7 are virtual sound sources spread out like a virtual surround sound system with virtual speakers surrounding the listener.

12. The method according to any of claims 1-11, wherein input signals are provided via two stereo input channels or as multiple discrete channels having signal processing already incorporated.

13. The method according to any of claims 2-12, wherein said method also comprises postprocessing including adding a delayed part of a transfer function as a virtual addition to the output sound in the headphones.

14. The method according to any of claims 1-13, wherein the method comprises using at least three direct sound sources, which are left, center and right (LCR), direct sound sources being such where the sound arrives directly from the sound source to someone’s ears, wherein the left and right are physical sound sources and the center channel is a virtual center channel providing a better tonal balance.

15. The method according to claim 14, wherein the other virtual sound sources are used to render the spatial sound field and the sound from these other virtual sound sources is delayed in relation to the direct sound from the LCR channels.

16. The method according to claim 14 or 15, wherein a direct sound output of said virtual center channel is produced by calculating the sum of the left and right physical sound source signals, preferably by performing left plus right (L+R) calculation on said signals.

17. The method according to claim 16, wherein difference signals, left minus right (L-R) and right minus left (R-L) of the left and right physical sound sources are amplified, preferably said signals are amplified so that all virtual sound source outputs summed together produce amplification of the indirect ambient sounds and reverberation by 6dB to 18dB, more preferably by 9dB to 14dB, and most preferably by 12dB in relation to their original derived level.

18. The method according to any of claims 1-17, wherein the digital sound processing unit is arranged to process sound on-line as it is being streamed in real time, for example while being recorded, broadcast, or played back, or process recorded sound off-line, for example computer files containing recoded audio that are processed and written back to disk, wherein said digital sound processing unit is a computer implemented unit or other type of software unit, a cloud server or other processing audio over the internet, ora combination thereof.