Artifact suppression in binaural audio
By employing time-domain processing techniques such as all-pass filters and comb filters with prime number-based delay parameters, the binaural Tenderer and re-renderer address the comb-filter artifact issue, improving the 3D sound immersion in binaural audio systems.
Patent Information
- Application Number
- PCT/US2024/057150
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-18
- Filing Date
- 2024-11-22
- Publication Date
- 2025-05-30
AI Technical Summary
Binaural audio systems suffer from artifacts such as the ITD-induced comb-filter effect, leading to corrupted headtracking and sound coloration, which degrade the 3D sound immersion experience.
The implementation of a binaural Tenderer and re-renderer that uses time-domain processing, including an all-pass filter and a bank of comb filters with delay parameters based on prime numbers, to decorrelate audio objects and apply head-related transform filters (HRTFs) while normalizing HRTFs to suppress the comb-filter effect.
This approach effectively suppresses the comb-filter effect, reducing computational complexity and latency, and is suitable for resource-constrained devices like mobile and wearable devices, thereby enhancing the 3D sound immersion experience.
Smart Images

Figure US2024057150_30052025_PF_FP_ABST
Abstract
Description
ARTIFACT SUPPRESSION IN BINAURAL AUDIO1. Cross-Reference to Related Applications
[0001] This application claims the benefit of International Application PCT / CN2023 / 133594 filed on 23 November 2023, U.S. Provisional Patent Application No. 63 / 612,292 filed on 19 December 2023, and U.S. Provisional Patent Application No. 63 / 661,512 filed on 18 June 2024, each of which is incorporated herein by reference in its entirety.2. Field of the Disclosure
[0002] Various example embodiments relate generally to binaural audio and, more specifically but not exclusively, to suppression of artifacts in binaural audio.3. Background
[0003] Binaural audio is a type and method of audio recording that reproduce the way humans naturally experience sound. More specifically, a binaural audio recording captures an audio scene under substantially the same conditions as someone listening to it, thereby achieving a very realistic 3D sound immersion effect. This end result may be accomplished, e.g., by setting up a specific recording environment in which a pair of microphones is spaced a head-width apart to replicate how human ears capture and process the incoming sound waves. In general, binaural audio can be generated by sound capture or through audio-signal processing. While there are several different ways to create convincing binaural recordings, a preferred way of reproducing them is through a pair of headphones. Devices that can reproduce binaural audio include, but are not limited to, mobile phones, laptop computers, wearable devices, such as head mounted display devices (e.g., extended reality (XR), augmented reality (AR), and virtual reality (VR) headsets), and various home or office devices.BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS
[0004] Disclosed herein are various examples, embodiments, features, and aspects of a method and system for generating binaural audio. In various examples, the binaural audio can be generated from channel-based audio or from another binaural audio. In one example, a method of generating binaural audio includes transforming input audio signals representing an audio scene intocorresponding decorrelated audio objects and generating respective left and right audio components by applying a head related transform filter (HRTF) to each of the decorrelated audio objects. The method also includes summing the left audio components and summing the right audio components to generate left and right output binaural signals, respectively. In some examples, the decorrelation of audio objects is performed using an all-pass filter serially connected with a bank of comb filters whose delay parameters are selected based on different respective prime numbers. In some examples, different HRTFs are applied to low- and high-frequency signal components. In some examples, the HRTFs are normalized with respect to an HRTF corresponding to a reference headrotation angle.
[0005] According to an example embodiment, provided is a method of generating binaural audio, the method comprising: transforming a set of first audio signals corresponding to an audio scene into a set of first audio objects; decorrelating the set of first audio objects to generate a plurality of first decorrelated audio objects; applying a first HRTF to each of the first decorrelated audio objects to generate a respective left audio component and a respective right audio component; adding the respective left audio components to generate a left output binaural signal; and adding the respective right audio components to generate a right output binaural signal.
[0006] According to another example embodiment, provided is an audio system for generating binaural audio, the audio system comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: transform a set of first audio signals corresponding to an audio scene into a set of first audio objects; decorrelate the set of first audio objects to generate a plurality of first decorrelated audio objects; apply a first HRTF to each of the first decorrelated audio objects to generate a respective left audio component and a respective right audio component; add the respective left audio components to generate a left output binaural signal; and add the respective right audio components to generate a right output binaural signal.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Other aspects, features, and benefits of various disclosed embodiments will become more fully apparent, by way of example, from the following detailed description and the accompanying drawings, in which:
[0008] FIG. 1 is a schematic diagram illustrating interaural differences of time and intensity for an ideal spherical head according to one example.
[0009] FIG. 2 is a block diagram illustrating a 2-to-3 upmixer according to one example.
[0010] FIG. 3 is a schematic diagram illustrating transformation of the audio objects generated with the 2-to-3 upmixer of FIG. 2 to binaural audio signals according to one example.
[0011] FIG. 4 graphically illustrates a comb-filter effect corresponding to the transformation illustrated in FIG. 3 according to one example.
[0012] FIGS. 5A-5B are schematic diagrams illustrating transformations of the audio objects generated with the 2-to-3 upmixer of FIG. 2 to binaural audio signals according to additional examples.
[0013] FIG. 6 graphically illustrates a comb-filter effect corresponding to the transformations illustrated in FIGS. 5A-5B according to one example.
[0014] FIG. 7 is a block diagram of a decorrelator according to one example.
[0015] FIGS. 8A-8C graphically illustrate an effect of the decorrelator of FIG. 7 on binaural audio signals according to one example.
[0016] FIG. 9 is a block diagram illustrating a decorrelation-amount control circuit used with the decorrelator of FIG. 7 according to some examples.
[0017] FIG. 10 graphically illustrates the magnitude response difference for sequential samples during headtracking rotations according to one example.
[0018] FIG. 11 is a block diagram illustrating a binaural Tenderer circuit employing a lowpass filter according to one example.
[0019] FIGS. 12A-12B graphically illustrate the effect of head related transform filter (HRTF) normalization used in a binaural re-renderer according to one example.
[0020] FIG. 13 is a block diagram illustrating a binaural Tenderer according to one example.
[0021] FIG. 14 is a block diagram illustrating a binaural re-renderer according to one example.
[0022] FIG. 15 is a block diagram illustrating a binaural Tenderer according to another example.
[0023] FIG. 16 is a block diagram illustrating a binaural re-renderer according to another example.
[0024] FIG. 17 is a block diagram illustrating a computing device one or more instances of which can be used in a binaural audio system according to various examples.
[0025] FIG. 18 is a flowchart illustrating a method of generating binaural audio according to various examples.DETAILED DESCRIPTION
[0026] FIG. 1 is a schematic diagram illustrating interaural differences of time and intensity for an ideal spherical head 102 subjected to sound waves 101 received from a distant sound source according to one example. In the example shown, the sound waves 101 arrive earlier in time at the right ear R than at the left ear L. This time-of- arrival difference is referred to as the interaural time difference (ITD). The sound intensity difference between the ipsilateral right ear R and the contralateral left ear L is referred to as the interaural level difference (ILD). The ILD is caused by the “shadowing” effect of the head 102, which prevents some of the incoming sound energy from reaching the contralateral left ear L.
[0027] In various examples, the ITD and ILD cues are effective in complementary ranges of frequencies of acoustic environments. More specifically, the ILDs are more pronounced at frequencies greater than approximately 1.5 kHz because the head 102 is large compared to the wavelengths of the incoming sound waves 101 at those frequencies, thereby producing substantial sound reflections. The ITDs, on the other hand, are relatively similar for all frequencies. However, for periodic sounds, the ITDs can be decoded unambiguously only at frequencies for which the maximum physically possible ITD is smaller than one half of the period of the waveform at that frequency. Since for a typical human head 102 the maximum possible ITD is about 660 ps, the ITDs are generally more useful at frequencies lower than about 1.5 kHz. For pure tones, theinteraural phase delay (IPD) can also be used due to the straightforward one-to-one mapping between the ITDs and the IPDs of pure tones.
[0028] In some cases, for a spherical head 102 having a uniform surface, the ITD, T, corresponding to the sound waves 101 can be approximated using Eq. (1) as follows: T = (2r / c) sin 0 (1) where r is the radius of the head 102; c is the speed of sound; and 0 is the azimuth angle. The approximation corresponding to Eq. ( 1 ) can be relatively accurate for frequencies lower than approximately 500 Hz but breaks down for higher frequencies. For frequencies higher than about2 kHz, the ITD corresponding to the sound waves 101 can be approximated using Eq. (2) as follows: T = r / c) 0 + sin 0) (2)In some examples, the radius r is about 9 cm.
[0029] The ILD values are not as well predicted by analytical models as the ITD values. In some examples, the ILDs can be obtained by applying suitable signal processing to microphone signals captured by two microphones corresponding to the left ear L and the right ear R, respectively. It has been found empirically that the ILDs tend to depend on the azimuth angle of the sound source, frequency, and distance from the sound sources to the head 102. For example, the ILDs produced by distant sound sources can become as large as 25 dB in magnitude at relatively high frequencies, and the ILDs can be even greater when the sound source is relatively close to one of the ears L and R.
[0030] Herein, a binaural Tenderer is a method, device, or module configured to generate binaural audio from channel-based audio. In some literature, a binaural Tenderer may be referred to as “virtualizer.” The binaural Tenderer is typically configured to mimic the shadow filter effect of the head, torso, and pinna. The output signals of the binaural Tenderer include the respective audio signals intended to be heard at the eardrum of each ear. The head-related filter is typically referred to as “head related transform filter” (HRTF). The basic characteristics of the HRTF include the above-mentioned ITD and ILD, which are used to provide cues for the head shadow effect.
[0031] Herein, a binaural re-renderer is a method, device, or module configured to generate a new binaural audio from an existing binaural audio, with the two binaural audios corresponding to two different azimuthal orientations of the head. In some examples, a binaural re-renderer operatesto decompose the existing (input) binaural audio into several virtual audio objects and then re-render those virtual audio objects with the HRTF corresponding to the changed azimuthal orientations of the head to generate the new (output) binaural audio. The usage of this technology is usually related to head-tracking, which is based on a stationary sound field that is independent of the head’s rotation.
[0032] In some examples, the binaural Tenderer and / or re-renderer suffer from an ITD-induced comb-filter effect. Example perceptible manifestations of the comb-filter effect include, but are not limited to, corrupted headtracking and / or sound coloration. Example embodiments disclosed herein provide a binaural Tenderer and re-renderer configured to suppress the comb-filter effect, thereby addressing the above-indicated problems in the state of the art. Some examples of such binaural Tenderer and re-renderer use time-domain processing, thereby providing solutions characterized by relatively low computational complexity and relatively small latency. The latter characteristics of these solutions provide a good fit, in terms of the memory, computational power allocation, and physical packaging (e.g., weight and size), to at least some of the most ubiquitous platforms, such as mobile devices, wearable devices, and edge-computing devices. For portable devices (especially wearable devices), which must be compact and lightweight, the reduced computational complexity of the proposed methods is particularly advantageous for increasing battery life and / or reducing device cost while keeping device weight within a reasonable envelope (e.g., via use of lesser performant processors with reduced cooling requirements, smaller batteries, etc.).Comb-Filter Effect
[0033] FIG. 2 is a block diagram illustrating a 2-to-3 upmixer 200 according to one example. Herein, an upmixer is a signal-processing module configured to up-mix channel-based audio to object-based audio having more audio objects than the number of input audio channels. In some examples, those audio objects are used as virtual speakers for a binaural Tenderer. The generation of “more” audio objects from the existing input channels typically includes separating a channel signal into two or more split signals that are then used to generate the corresponding audio objects.
[0034] In the example shown, the upmixer 200 takes two input channels, denoted 202L and 202R, respectively, and generates three output audio objects, denoted 222L, 222C, and 222R, respectively. The upmixer 200 includes first and second channel separation blocks 210i and 2102.The first channel separation block 210i operates to separate the input channel 202L into two parts, denoted as the “left-only” channel part Lo and the “left-center” channel part Lc. The second channel separation block 2102 similarly operates to separate the input channel 202R into two parts, denoted as the “right-only” channel part Ro and the “right-center” channel part Rc. Then, the first output audio object 222L is generated based on the channel part Lo. The second output audio object 222C is generated based on a combination of the channel parts Lc and Rc. The third output audio object 222R is generated based on the channel part Ro.
[0035] FIG. 3 is a schematic diagram illustrating a rendering 300 of the audio objects 222L, 222C, and 222R to binaural audio signals 304, 306 according to one example. For illustration purposes, only a part of the rendering 300 that is relevant to the explanation of the comb-filter effect is explicitly shown in FIG. 3. The shown part includes the rendering of the audio objects 222L and 222C to the left binaural signal 304 via the application of an HRTF 302.
[0036] The channel separation operation performed in the first channel separation block 21 Or of the upmixer 200 is not typically completely clean, which causes a residual correlation component Res to be present in each of the audio objects 222L and 222C as indicated in FIG. 3. In at least some examples, the residual correlation component Res manifests itself via the comb-filter effect. The comb-filter effect arises primarily due to the different respective ITDs being applied to the residual correlation component Res of the audio objects 222L and 222C as indicated in FIG. 3 by the corresponding delay values DelayO and Delay 1 (also see Eqs. (1 )-(2)).
[0037] Let us denote the residual correlation component Res as xrand ignore the ILD for the sake of simplicity. Then, the combined audio signal yr[n] arriving from the audio objects 222L and 222C to contribute to the left binaural signal 304 can be expressed as follows: yr[n] = xr[n] + xr[n — Ad] (3) where n is the time index; and d is the time difference between the delay values DelayO and Delay 1 expressed in the units of the time index. A z-domain representation of Eq. (3) is the transfer function H(z) expressed as follows:H(z) = 1 + z“Ad(4)A frequency-domain representation of the transfer function is obtained by the substitution z = e]U>, which produces the following expression:The magnitude response obtained from Eq. (5) is as follows:|H(e")| = V( + cosroAd)2+ (sinoiAd)2= 2 + 2cosa> d (6)
[0038] FIG. 4 graphically illustrates a comb-filter effect corresponding to the residual correlation component Res of the rendering 300 according to one example. More specifically, a magnitude response curve 402 shown in FIG. 4 is a plot of Eq. (6) corresponding to Ad = 26. As evident from FIG. 4, the magnitude response curve 402 has a plurality of periodic peaks and valleys in the frequency domain. The corresponding signal modulation will disadvantageously impose a “coloring” on the residual correlation component Res in the left binaural audio signal 304. A person of ordinary skill in the art will readily understand that such coloring may also be present in the right binaural audio signal 306.
[0039] FIGS. 5A-5B are schematic diagrams illustrating renderings 5OOo, 50030 of the audio objects 222E, 222C, and 222R to respective binaural audio signals 304, 504 according to another example. The renderings 5OOo, 50030 correspond to a head-tracking Tenderer capable of changing the angular orientation of the HRTF 302. More specifically, the rendering 5OOo illustrated in FIG. 5A corresponds to the angular orientation of the HRTF 302 represented by an orientation vector 501. The rendering 50030 similarly illustrated in FIG. 5B corresponds to the angular orientation of the HRTF 302 represented by an orientation vector 502. The angle between the orientation vectors 501 and 502 is 30 degrees. In general, the head-tracking Tenderer will take the HRTF orientation vector as an input and render the binaural audio signals from the audio objects accordingly.
[0040] FIG. 6 graphically illustrates a comb-filter effect corresponding to the residual correlation component Res of the renderings 5OOo, 50030 according to one example. The magnitude response curve 402 shown in FIG. 6 corresponds to the rendering 5OOo (FIG. 5A) and is the same as the magnitude response curve 402 of FIG. 4. As already indicated above, the magnitude response curve 402 corresponds to Ad = 26. A magnitude response curve 602 shown in FIG. 6 corresponds to the rendering 50030 (FIG. 5B) for which Ad = 27. Note that, over the frequency range shown, the magnitude response curve 602 has one more peak than the magnitude response curve 402. As a result, for relatively high normalized frequencies (e.g., >0.6), the peaks of the magnitude response curve 602 are approximately aligned with the valleys of the magnitude response curve 402, and viceversa. This means that, for some head rotations, the modulation magnitude of some specific high frequencies can be very large, e.g., larger than 20 dB. This large modulation magnitude is sometimes referred to as the comb-filter interlacing artifact. This artifact can disadvantageously corrupt the head-tracking effect and significantly degrade the 3D sound immersion experience for the user.Decorrelator
[0041] FIG. 7 is a block diagram of a decorrelator 700 according to one example. The decorrelator 700 can be used, e.g., to significantly reduce or substantially eliminate residual correlation components among up-mixed audio objects, such as the above-described audio objects 222L, 222C, and 222R. For example, before the HRTF 302, each audio object can be subjected to processing in the decorrelator 700. In the example shown, the decorrelator 700 operates in the time domain, which enables the corresponding Tenderer or re-renderer to maintain relatively low latency and have a relatively low computational complexity.
[0042] The decorrelator 700 includes an all-pass filter 710 and a plurality of comb filters 72Oo- 720N connected in parallel to one another and in series with the all-pass filter 710. In some examples, the number N is N=3. In other examples, other values of N can also be used. The gain(s) and delay of each individual filter is adjustable. An important idea informing the design of the decorrelator 700 is to impose substantially random phase shifts onto an input signal 702 as explained in more detail below. Effectively, the decorrelator 700 smears transients, thereby making them less localized. As long as the smearing occurs on a time scale shorter than about 5-10 ms, the smearing does not produce additional perceptual artifacts for humans while largely disrupting the constructive / destructive interference of the residual correlation components emanating from different audio objects.
[0043] The amplitude response of the all-pass filter 710 is “1” at each frequency while its phase response (which determines the delay as a function of frequency) can be selected. In one example, the all-pass filter 710 is constructed by cascading a feedback comb-filter with a feedforward combfilter in a structure known as the Schroeder all-pass section. The corresponding difference equation describing the all-pass filter 710 is as follows:where x is the input signal 702; yais an output signal 712; n is the time index; dais a delay parameter; and gais the gain parameter. The z-domain transfer function H(z) of the all-pass filter 710 is expressed as follows:The frequency response of the all-pass filter 710 is obtained by the substitution z = e]U>in Eq. (8).
[0044] The parallel comb filters 72OO-72ON have different respective delay parameters di, which are selected to be larger than da. This selection aims to increase the density of peaks and valleys in the individual filter’s transfer functions. The corresponding difference equation describing the comb filter 720i is as follows:where yais the output signal 712 of the all-pass filter 710, which serves as an input signal to each of comb filters 720i; yct is the output signal of the comb filter 720i; and gt land gti2are the gain parameters. An adder 730 operates to sum the output signals of different comb filters 720i to generate an output signal 732 (which is denoted as yc) of the decorrelator 700. The corresponding mathematical expression for the output signal 732 is as follows: yc[n] = s =o{yc,i[n]} (10)
[0045] In some examples, the delay parameters di are selected from prime numbers. Larger delay parameters typically introduce stronger reverberation effects, whereas smaller delay parameters typically introduce a stronger metallic distortion, both which are generally undesirable. As such, a trade-off selection that weights and balances these considerations may be preferred. As an illustration, Table 1 provides a set of delay and gain values for several comb filters 720i according to one example.Table 1: Example parameters for comb filters 720i of decorrelator 700.
[0046] FIGS. 8A-8C graphically illustrate an effect of the decorrelator 700 according to one example. More specifically, FIG. 8A shows time-dependent spectrograms 802, 804 of stereo whitenoise signals 202L, 202R. These stereo white-noise signals 202L, 202R are applied to the upmixer 200 to generate the corresponding audio objects 222L, 222C, and 222R. These audio objects 222L, 222C, and 222R are then subjected to the HRTF 302 to generate the corresponding binaural audio signals 304, 306. Time-dependent spectrograms 814, 816 and 824, 826 of those binaural audio signals 304, 306 are illustrated in FIGS. 8B and 8C, respectively. FIG. 8B represents the case in which the decorrelator 700 is not used. FIG. 8C similarly represents the case in which the decorrelator 700 is used. The virtual speaker angle is 60 degrees in both cases. During the time interval represented in each of FIGS. 8B and 8C, the HRTF orientation vector changes its orientation from 89 degrees to 90 degrees (also see FIGS. 5A-5B).
[0047] The time-dependent spectrogram 816 shown in FIG. 8B clearly manifests the presence of the comb-filter interlacing artifact on the ipsilateral ear due to the HRTF orientation vector rotation from 89 degrees to 90 degrees when the decorrelator 700 is not used. When the decorrelator 700 is individually applied to the corresponding audio objects 222L, 222C, and 222R in the same rotation of the HRTF 302, the comb-filter interlacing artifact is substantially eliminated, which is clearly seen from the comparison of the time-dependent spectrograms 816 (FIG. 8B) and 826 (FIG. 8C). However, the time-dependent spectrograms 814 and 826 (FIG. 8C) still exhibit some reverberation effects and signal coloring. To address the latter effects, the amount of decorrelation may need to be controlled and / or optimized as further illustrated in FIG. 9.
[0048] FIG. 9 is a block diagram illustrating a decorrelation-amount control circuit 900 used with the decorrelator 700 according to some examples. In addition to the decorrelator 700, the circuit 900 includes weighting modules 902, 906 and an adder 910. The weighting module 902 is configured to apply a decorrelation factor g to the output signal 732 generated by the decorrelator 700 as described above. The value of the decorrelation factor g is selectable from the interval [0, 1].The weighting module 906 is configured to apply a weighting factor (1-g) to the input signal 702 of the decorrelator 700. The adder 910 is configured to generate an output signal 912 of the circuit 900 by summing the weighted signals generated with the weighting modules 902, 906. When the decorrelation factor g is selected to be g=0, the output signal 912 is the same as the input signal 702 of the decorrelator 700. When the decorrelation factor g is selected to be g=l, the output signal 912 is the same as the output signal 732 of the decorrelator 700. When an intermediate value from the interval [0,1] is selected for the decorrelation factor g, the output signal 912 is a corresponding mixture of the signals 702 and 732.Lowpass Filtering
[0049] The above-mentioned metallic distortion and unwanted reverberation artifacts warrant additional considerations for the design of a decorrelator. In particular, it should be noted that timedomain implementations of the decorrelator have inherent limitations with respect to fully resolving those artifacts. When circuit 900 is configured to use a value of the decorrelation factor g that is smaller than one, the output signal 912 will contain a residual correlation component. This component is going to be less noticeable when no headtracking takes place. However, in a headtracking use case, the azimuth of the user head is going to change in real-time, which causes the time difference of a pair of up-mixed audio objects to be exposed to block-by -block changes. The comb-filter interlacing artifact is especially prominent around +90°, e.g., because the time difference tends to be relatively large in that angular region. For example, a small rotation of the head can bring the above-explained large change of the magnitude response in the high frequency range into play, which triggers the comb teeth interlacing artifact.
[0050] In typical applications, the human head-rotation speed and the sensor sampling rate are such that the relative rotation between two consecutive samples is relatively small. Denoting the sensor sampling rate as fsQand the maximum head rotation speed asthe maximum angle interval can be expressed as V fsQ. The corresponding delay time difference value, dmflx, measured in samples can be expressed as follows: dmax= max (|T(cr0) - T(a0± / so)l) / s (H) Where the function T(cr) is given by Eq. (12):
[0051] FIG. 10 graphically illustrates the magnitude response difference (MRD) for sequential samples during head rotations according to one example. More specifically, curves 1002 and 1004 are plots of Eq. (11) corresponding to the following parameter values:= 200 degree / s; fs0=20 Hz; head radius r = 0.09 m; the sound speed c = 343 m / s; the max value is around a0=andthe max sample interval between two renders is about 2 samples. The curve 1002 represents the MRD between two sequential samples separated by one sampling interval. The curve 1004 similarly represents the MRD between two sequential samples separated by two sampling intervals. It should be noted that the curve 1004 manifests a more-pronounced interlacing problem than the curve 1002.
[0052] The graphs shown in FIG. 10 indicate that lower frequencies suffer less from the interlacing artifact during head rotation than higher frequencies. In addition, the time difference will cause the phase difference to wrap around 2n for higher frequencies. From the phase wrapping perspective, it may be beneficial to restrict the ITD application to lower frequencies in at least some cases.
[0053] FIG. 11 is a block diagram illustrating a binaural Tenderer circuit 1100 employing a lowpass filter 1110 according to one example. More specifically, the lowpass filter 1110 is configured to split an input signal 1102 into a low frequency (EF) component 1112 and a high frequency (HF) component 114. In one example, the lowpass filter 1110 is implemented using an 8- tap Butterworth lowpass filter whose cutoff frequency f0is selected to be f0= 2 kHz. In other examples, other cutoff frequencies for the lowpass filter 1110 can also be selected.
[0054] The LF component 1112 is sequentially subjected to ITD processing in an ITD module 1120 and to ILD processing in an ILD module 1130i . In contrast, the HF component 1114 is only subjected to ILD processing in an ILD module 11302. An adder 1140 is configured to generate an output signal 1142 of the binaural Tenderer circuit 1100 by summing the output signals of the ILD modules 1130i and 11302.Binaural Re-rendering
[0055] For stereo or channel-based audio, a binaural Tenderer can improve the localization, externalization, and spatial characteristics of the delivered audio. However, for a binaural input, the same binaural Tenderer will impose double processing. In at least some examples, such double processing may cause at least some characteristics of the delivered audio, such as the above- mentioned localization and other important features, to be undesirably degraded. In many cases, after such double processing, the binaural content is no longer in accordance with the original creative intent.
[0056] To address these problems and also to provide support for head-tracking of binaural input, some embodiments disclosed herein employ a re-rendering process with HRTF normalization. The normalization coefficient can be calculated, e.g., based on the HRTF of the original (reference) state where no rotation has occurred yet. Then, every HRTF applied thereafter is normalized with coefficients such that the HRTF applied is the difference between the angle’s HRTF and the original HRTF.
[0057] Let a binaural process be in accordance with the following expressions:where i is the object number; h is the magnitude response; and cp is the time-delay-caused phase shift. The index I represents the left ear, and the index r represents the right ear. Let the HRTF parameters calculated from the head rotation angle 0 be denoted as h(((0) and <p(((0). Then, a normalized re-render process can be expressed as follows:y ir (o)1 L JThis process will flatten the magnitude response and reset the time delay at 0 degree to preserve the input binaural effect. For example, a sphere model Tenderer can be described as: y[n] = bQx[n — d] + b-jx n — d — 1] + aoy[n — 1] (17)The corresponding binaural re-renderer is then expressed as follows:
[0058] FIGS. 12A-12B graphically illustrate the effect of HRTF normalization used in a binaural re-renderer according to one example. In the example shown, a virtual object is positioned at an angle of 60 degrees. Different panels in FIG. 12 graphically illustrate the magnitude response of the HRTF to the ipsilateral ear for different respective azimuth head-rotation angles before and after normalization computed in accordance with Eqs. (17)-(18). Inspection of the various shown magnitude response curves indicates that the HRTF normalization beneficially flattens the magnitude response at zero degree while accounting for the difference for other angles.System Design
[0059] FIG. 13 is a block diagram illustrating a binaural Tenderer 1300 according to one example. The Tenderer 1300 is configured to convert a channel-based audio input into a binaural audio output. In the example shown, the channel-based audio input includes n channels 13021 - 1302n, where n is an integer greater than one. For a conventional stereo input, the number n is n=2. The binaural audio output includes a left binaural signal 1362L and a right binaural signal 1362R.
[0060] The Tenderer 1300 includes an n-to-N upmixer 1310 configured to convert the channelbased audio input 13021 - 1302ninto a corresponding plurality of audio objects 1322I-1322N. In some examples, the number N is N=3. In some examples, the n-to-N upmixer 1310 can be implemented using the above described 2-to-3 upmixer 200. In such examples, the corresponding up-mixing process can be similar to that described above in reference to FIG. 2.
[0061] The Tenderer 1300 also includes a plurality of decorrelators 900I-900N. Each of the decorrelators 900I-900N is a respective instance of the decorrelator 900 described above in reference to FIG. 9. Each of the decorrelators 900I-900N operates to decorrelate the respective one of the audio objects 1322I-1322N, thereby producing a respective one of decorrelated objects 912I-912N. As already explained above, the value of the respective decorrelation factor g for each of the decorrelators 900I-900N is independently controllable.
[0062] The Tenderer 1300 also includes a plurality of HRTF circuits 1340I-1340N. Each of the HRTF circuits 1340I-1340N is constructed based on the circuit architecture described in reference to the circuit 1100 (FIG. 11). One difference between the circuits 1100 and 1340i is that the latter circuit is configured to generate components for both the left and right binaural signals 1362L, 1362R. Accordingly, the circuit 1340i has two instances of the ITD module 1120, which aredenoted 1120L and 1120R, respectively. The ITD module 1120L is configured to apply the ITD corresponding to the left ear. The ITD module 1120R is similarly configured to apply the ITD corresponding to the right ear. The circuit 1340i also has two instances of an ILD module 1350, which are denoted 1350L and 1350R, respectively. Each ILD module 1350 includes a respective circuit set including the ILD modules 1130i and 11302 and the adder 1140 connected as indicated in FIG. 11. The ILD module 1350L is configured to apply the ILD corresponding to the left ear. The ILD module 1350R is similarly configured to apply the ILD corresponding to the right ear. Output signals of the circuit 1340i are denoted in FIG. 13 using the reference numerals 1352L; and 1352Ri, where i=l, 2, ..., N. The signal 1352L; provides a component for the left binaural signal 1362L. The signal 1352Ri provides a component for the right binaural signal 1362R.
[0063] The Tenderer 1300 also includes adders 1360L and 1360R. The adder 1360L is configured to generate the left binaural signal 1362L by summing the signals 1352LI-1352LN received from the HRTF circuits 1340I-1340N. The adder 1360R is similarly configured to generate the right binaural signal 1362R by summing the signals 1352RI-1352RN received from the HRTF circuits 1340I-1340N.
[0064] FIG. 14 is a block diagram illustrating a binaural re-renderer 1400 according to one example. The binaural re-renderer 1400 is generally similar to the binaural Tenderer 1300 (FIG. 13) but has several modifications, which are described in more detail below. The description of the unmodified elements is already given above in reference to FIG. 13.
[0065] The re-renderer 1400 is configured to convert a binaural audio input into a binaural audio output. In the example shown, the binaural audio input includes a left binaural signal 1402L and a right binaural signal 1402R. The binaural audio output includes a left binaural signal 1462L and a right binaural signal 1462R.
[0066] In the re-renderer 1400, the upmixer 1310 is a 2-to-N upmixer. Another difference between the Tenderer 1300 and the re-renderer 1400 is that, in the latter, HRTF circuits 1440I-1440N are used in place of the HRTF circuits 1340I-1340N. An HRTF circuit 1440i differs from a corresponding HRTF circuit 1340i in that the HRTF circuit 1440i implements the above-described HRTF normalization. The corresponding circuit modification involves: (i) replacing the ITD modules 1120L and 1120R by normalized ITD modules 1420L and 1420R, respectively, and (ii)replacing the ILD modules 1350L and 1350R by normalized ILD modules 1450L and 1440R, respectively. Output signals of the circuit 1440i are denoted in FIG. 14 using the reference numerals 1452Li and 1452Ri, where i=l, 2, ..., N. The adder 1360L is configured to generate the left binaural signal 1462L by summing the signals 1452LI-1452LN received from the HRTF circuits 14401- 1440N- The adder 1360R is similarly configured to generate the right binaural signal 1462R by summing the signals 1452RI-1452RN received from the HRTF circuits 1440I-1440N.
[0067] FIG. 15 is a block diagram illustrating a binaural Tenderer 1500 according to another example. The Tenderer 1500 is configured to convert a channel-based audio input into a binaural audio output. In the example shown, the channel-based audio input includes n channels 15021 - 1502n, where n is an integer greater than one. For a conventional stereo input, the number n is n=2. The binaural audio output includes a left binaural signal 1562L and a right binaural signal 1562R.
[0068] The Tenderer 1500 includes a multichannel lowpass filter 1510 configured to separate each of the channel signals 1502i-1502ninto a respective low-frequency component 1512j and a respective high-frequency component 1513j, where j = 1, 2, ..., n. With respect to each individual channel signal 1502j, the function of the multichannel lowpass filter 1510 is analogous to that of the above-described lowpass filter 1110 (FIG. 11).
[0069] The Tenderer 1500 includes n-to-N upmixers 1310i and 13102. The upmixer 1310i operates in the low-frequency range and is configured to convert the low-frequency components 1512i- 1512ninto a corresponding plurality of low-frequency audio objects 1522I-1522N. The upmixer 13102 operates in the high-frequency range and is configured to convert the high-frequency components 15131-1513ninto a corresponding plurality of high-frequency audio objects 15231- 1523N.
[0070] The Tenderer 1500 also includes a plurality of decorrelators 9001-9002N. Each of the decorrelators 9001-9002N is a respective instance of the decorrelator 900 described above in reference to FIG. 9. Each of the decorrelators 900I-900N operates to decorrelate the respective one of the audio objects 1522I-1522N, thereby producing a respective one of decorrelated objects 912i - 912N- Each of the decorrelators 900N+I-9002N similarly operates to decorrelate the respective one of the audio objects 1523I-1523N, thereby producing a respective one of decorrelated objects 912N+I-9122N. The value of the respective decorrelation factor g for each of the decorrelators 900I-9002N is independently controllable.
[0071] The Tenderer 1500 also includes a plurality of HRTF circuits 1540I-1540N. Each of the HRTF circuits 1540I-1540N is constructed based on the circuit architecture of the low-frequency branch of the circuit 1100 described above in reference to FIG. 11. An individual circuit 1540i is configured to process the decorrelated object 912i to generate components for both the left and right binaural signals 1562E, 1562R. As such, circuit 1540i has first and second parallel branches, each including serially connected modules 1120 and 1130i . The first branch is configured to generate a low-frequency component 1542E; for the left binaural signal 1562E and includes the ITD module 1120E and the IED module 1130IL. The second branch is configured to generate a low-frequency component 1542R; for the right binaural signal 1562R and includes the ITD module 1120R and the IED module 1130IR.
[0072] The Tenderer 1500 also includes a plurality of HRTF circuits 1550I-1550N. Each of the HRTF circuits 1550I-1550N is constructed based on the circuit architecture of the high-frequency branch of the circuit 1100 described above in reference to FIG. 11. An individual circuit 1550i is configured to process the decorrelated object 912N-H to generate components for both the left and right binaural signals 1562L, 1562R. As such, circuit 1550i has first and second parallel branches, each including a respective instance of the ILD module 11302. The first branch is configured to generate a high-frequency component 1552L; for the left binaural signal 1562L. The second branch is configured to generate a high-frequency component 1552R; for the right binaural signal 1562R.
[0073] The Tenderer 1500 also includes adders 1560E and 1560R. The adder 1560E is configured to generate the left binaural signal 1362L by summing the signals 1542LI-1542LN received from the HRTF circuits 1540I-1540N and the signals 1552LI-1552LN received from the HRTF circuits 1550I-1550N. The adder 1560R is similarly configured to generate the right binaural signal 1562R by summing the signals 1542RI-1542RN received from the HRTF circuits 15401- 1540N and the signals 1552RI- 1552RN received from the HRTF circuits 1550I-1550N.
[0074] FIG. 16 is a block diagram illustrating a binaural re-renderer 1600 according to one example. The binaural re-renderer 1600 is generally similar to the binaural Tenderer 1500 (FIG. 15)but has several modifications, which are described in more detail below. The description of the unmodified elements is already given above in reference to FIG. 15.
[0075] The re-renderer 1600 is configured to convert a binaural audio input into a binaural audio output. In the example shown, the binaural audio input includes a left binaural signal 1602L and a right binaural signal 1602R. The binaural audio output includes a left binaural signal 1662L and a right binaural signal 1662R.
[0076] In the re-renderer 1600, each of the upmixers 1310i and 1310i is a 2-to-N upmixer.Another difference between the Tenderer 1500 and the re-renderer 1600 is that, in the latter, HRTF circuits 1640I-1640N are used in place of the HRTF circuits 1540I-1540N, and HRTF circuits 16501 - 1650N are used in place of the HRTF circuits 1550I-1550N. An HRTF circuit 1640i differs from a corresponding HRTF circuit 1540i in that the HRTF circuit 1640i implements the above-described HRTF normalization. The corresponding circuit modification involves: (i) replacing the ITD modules 1120L and 1120R by normalized ITD modules 1620L and 1620R, respectively, and (ii) replacing the ILD modules 1130IL and 1130IR by normalized ILD modules 1630IL and 1630IR, respectively. Output signals of the circuit 1640i are denoted in FIG. 16 using the reference numerals 1642Li and 1642Ri, where i=l, 2, ..., N. An HRTF circuit 1650i differs from a corresponding HRTF circuit 1550i in that the HRTF circuit 1650i implements the above-described HRTF normalization. The corresponding circuit modification involves replacing the ILD modules 11302L and 11302R by normalized ILD modules 16302L and 16302R, respectively. Output signals of the circuit 1650i are denoted in FIG. 16 using the reference numerals 1652L; and 1652Ri,
[0077] The adder 1560L is configured to generate the left binaural signal 1662L by summing the signals 1642LI-1642LN received from the HRTF circuits 1640I-1640N and the signals 1652L1- 1652LN received from the HRTF circuits 1650I-1650N. The adder 1560R is similarly configured to generate the right binaural signal 1662R by summing the signals 1642RI-1642RN received from the HRTF circuits 1640I-1640N and the signals 1652RI-1652RN received from the HRTF circuits 16501- 1650N.Example Hardware
[0078] FIG. 17 is a block diagram of an example computing device 1700 according to various examples. In some examples, the computing device 1700 is configured to various Tenderers, re- renderers, circuits, and devices disclosed herein.
[0079] The computing device 1700 of FIG. 17 is illustrated as having a number of components, but any one or more of these components may be omitted or duplicated, as suitable for the application and setting. In some embodiments, some or all of the components included in the computing device 1700 may be attached to one or more motherboards and enclosed in a housing. In some embodiments, some of those components may be fabricated onto a single system-on-a-chip (SoC) (e.g., the SoC may include one or more electronic processing devices 1702 and one or more storage devices 1704). Additionally, in various embodiments, the computing device 1700 may not include one or more of the components illustrated in FIG. 17, but may include interface circuitry for coupling to the one or more components using any suitable interface (e.g., a Universal Serial Bus (USB) interface, a High-Definition Multimedia Interface (HDMI) interface, a Controller Area Network (CAN) interface, a Serial Peripheral Interface (SPI) interface, an Ethernet interface, a wireless interface, or any other appropriate interface). For example, the computing device 1700 may not include a display device 1710, but may include display device interface circuitry (e.g., a connector and driver circuitry) to which an external display device 1710 may be coupled.
[0080] The computing device 1700 includes a processing device 1702 (e.g., one or more processing devices). As used herein, the terms “electronic processor device” and “processing device” interchangeably refer to any device or portion of a device that processes electronic data from registers and / or memory to transform that electronic data into other electronic data that may be stored in registers and / or memory. In various embodiments, the processing device 1702 may include one or more digital signal processors (DSPs), application-specific integrated circuits (ASICs), central processing units (CPUs), graphics processing units (GPUs), server processors, or any other suitable processing devices.
[0081] The computing device 1700 also includes a storage device 1704 (e.g., one or more storage devices). In various embodiments, the storage device 1704 may include one or more memory devices, such as random-access memory (RAM) devices (e.g., static RAM (SRAM) devices, magnetic RAM (MRAM) devices, dynamic RAM (DRAM) devices, resistive RAM (RRAM) devices, or conductive-bridging RAM (CBRAM) devices), hard drive-based memorydevices, solid-state memory devices, networked drives, cloud drives, or any combination of memory devices. In some embodiments, the storage device 1704 may include memory that shares a die with the processing device 1702. In such an embodiment, the memory may be used as cache memory and include embedded dynamic random-access memory (eDRAM) or spin transfer torque magnetic random-access memory (STT-MRAM), for example. In some embodiments, the storage device 1704 may include non-transitory computer readable media having instructions thereon that, when executed by one or more processing devices (e.g., the processing device 1702), cause the computing device 1700 to perform any appropriate ones of the methods disclosed herein below or portions of such methods.
[0082] The computing device 1700 further includes an interface device 1706 (e.g., one or more interface devices 1706). In various embodiments, the interface device 1706 may include one or more communication chips, connectors, and / or other hardware and software to govern communications between the computing device 1700 and other computing devices. For example, the interface device 1706 may include circuitry for managing wireless communications for the transfer of data to and from the computing device 1700. The term “wireless” and its derivatives may be used to describe circuits, devices, systems, methods, techniques, communications channels, etc., that may communicate data via modulated electromagnetic radiation through a nonsolid medium. The term does not imply that the associated devices do not contain any wires, although in some embodiments they might not. Circuitry included in the interface device 1706 for managing wireless communications may implement any of a number of wireless standards or protocols, including but not limited to Institute for Electrical and Electronic Engineers (IEEE) standards including Wi-Fi (IEEE 802.11 family), IEEE 802.16 standards, Long-Term Evolution (LTE) project along with any amendments, updates, and / or revisions (e.g., advanced LTE project, ultramobile broadband (UMB) project (also referred to as “3GPP2”), etc.). In some embodiments, circuitry included in the interface device 1706 for managing wireless communications may operate in accordance with a Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High Speed Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE network. In some embodiments, circuitry included in the interface device 1706 for managing wireless communications may operate in accordance with Enhanced Data for GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). In someembodiments, circuitry included in the interface device 1706 for managing wireless communications may operate in accordance with Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolution-Data Optimized (EV-DO), and derivatives thereof, as well as any other wireless protocols that are designated as 3G, 4G, 5G, and beyond. In some embodiments, the interface device 1706 may include one or more antennas (e.g., one or more antenna arrays) configured to receive and / or transmit wireless signals.
[0083] In some embodiments, the interface device 1706 may include circuitry for managing wired communications, such as electrical, optical, or any other suitable communication protocols. For example, the interface device 1706 may include circuitry to support communications in accordance with Ethernet technologies. In some embodiments, the interface device 1706 may support both wireless and wired communication, and / or may support multiple wired communication protocols and / or multiple wireless communication protocols. For example, a first set of circuitry of the interface device 1706 may be dedicated to shorter-range wireless communications such as Wi-Fi or Bluetooth, and a second set of circuitry of the interface device 1706 may be dedicated to longer- range wireless communications such as global positioning system (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, or others. In some other embodiments, a first set of circuitry of the interface device 1706 may be dedicated to wireless communications, and a second set of circuitry of the interface device 1706 may be dedicated to wired communications.
[0084] The computing device 1700 also includes battery / power circuitry 1708. In various embodiments, the battery / power circuitry 1708 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuitry for coupling components of the computing device 1700 to an energy source separate from the computing device 1700 (e.g., to AC line power).
[0085] The computing device 1700 also includes a display device 1710 (e.g., one or multiple individual display devices). In various embodiments, the display device 1710 may include any visual indicators, such as a heads-up display, a computer monitor, a projector, a touchscreen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.
[0086] The computing device 1700 also includes additional input / output (I / O) devices 1712. In various embodiments, the TO devices 1712 may include one or more data / signal transfer interfaces,audio I / O devices (e.g., microphones or microphone arrays, speakers, headsets, earbuds, alarms, etc.), audio codecs, video codecs, printers, sensors (e.g., thermocouples or other temperature sensors, humidity sensors, pressure sensors, vibration sensors, etc.), image capture devices (e.g., one or more cameras), human interface devices (e.g., keyboards, cursor control devices, such as a mouse, a stylus, a trackball, or a touchpad), etc. In some examples, the I / O devices 1712 include audio devices capable of implementing various binaural techniques described above, such as XR / AR / VR headsets and earbuds or other suitable XR / AR / VR audio devices.
[0087] Depending on the specific embodiment, various components of the interface devices 1706 and / or I / O devices 1712 can be configured to output suitable control signals, receive suitable control / telemetry signals, and receive and transmit data streams. In some examples, the interface devices 1706 and / or I / O devices 1712 include one or more analog-to-digital converters (ADCs) for transforming received analog signals into a digital form suitable for operations performed by the processing device 1702 and / or the storage device 1704. In some additional examples, the interface devices 1706 and / or I / O devices 1712 include one or more digital-to-analog converters (DACs) for transforming digital signals provided by the processing device 1702 and / or the storage device 1704 into an analog form suitable for being transmitted through a communication channel.System-Implemented Method of Generating Binaural Audio
[0088] FIG. 18 is a flowchart illustrating a method 1800 of generating binaural audio according to some examples. Various embodiments of the method 1800 can be implemented using various systems, devices, and hardware described above, e.g., in reference to FIGS. 13-17.
[0089] At a block 1802 “Transform input audio signals into corresponding audio objects,” the method 1800 includes transforming a set of first audio signals corresponding to an audio scene into a set of first audio objects. In some examples, the set of first audio signals includes the channels 1302i-1302n(FIG. 13), the signals 1402L and 1402R (FIG. 14), the channels 15021- 1502n(FIG. 15), or the signals 1602L and 1602R (FIG. 16). The transforming operations of the block 1802 can be performed using one or more upmixers 1310, e.g., as illustrated in FIGS. 13-16. The set of first audio objects may include the audio objects 1322I-1322N, 1522I-1522N, and / or 1523I-1523N. In some examples, the number of audio objects in the set of first audio objects is greater than the number of signals in the set of first audio signals.
[0090] At a block 1804 “Decorrelate audio objects,” the method 1800 includes decorrelating the set of first audio objects to generate a plurality of first decorrelated audio objects. The decorrelating operations of the block 1804 can be performed using one or more decorrelators 700 and / or 900 (FIGS. 7, 9). In some examples, the decorrelating operations include: (i) applying an individual one of the first audio objects to an all-pass filter (e.g., 710, FIG. 7) to generate a respective first filtered signal (e.g., 712, FIG. 7); (ii) applying the respective first filtered signal to a plurality of comb filters (e.g., 72OO-72ON, FIG. 7) to generate a plurality of respective second filtered signals; and (iii) summing (e.g., with the adder 730, FIG. 7) the plurality of respective second filtered signals to generate a respective summed signal (e.g., 732, FIG. 7). In some examples, wherein the all-pass filter comprises a Schroeder all-pass section. In some examples, different ones of the comb filters have different respective delay parameters selected based on different respective prime numbers. In some examples, the respective summed signal represents a respective one of the first decorrelated audio objects. In some examples, operations of the block 1804 also include generating a respective one of the first decorrelated audio objects using a weighted sum of the respective summed signal and the individual one of the first audio objects.
[0091] At a block 1806 “Apply HRTF to each of decorrelated audio objects to generate respective left and right audio components,” the method 1800 includes applying a first head related transform filter (HRTF) to each of the first decorrelated audio objects to generate a respective left audio component and a respective right audio component. An example HRTF that can be used in the block 1806 is the HRTF 302 described in reference to FIG. 3. In some examples, operations of the block 1806 include performing headtracking via rotation of the first HRTF. In some examples, the operations of the block 1806 include: (i) splitting an individual one of the first decorrelated audio objects into a respective low-frequency audio object and a respective high-frequency audio object; (ii) applying interaural-time-difference (ITD) processing and interaural-level-difference (ILD) processing to the respective low-frequency audio object; and (iii) applying ILD processing but no ITD processing to the respective high-frequency audio object. In some examples, the first HRTF is normalized with respect to an HRTF corresponding to a reference head-rotation angle. The normalization operations can be performed using one or more sets of the modules 1420, 1450, 1620, and / or 1630 (FIGS. 14, 16).
[0092] At a block 1808 “Add respective left and right audio components to generate output binaural signals and optionally direct output binaural signals to audio-rendering equipment,” the method 1800 includes adding the respective left and right audio components to generate left and right output binaural signals. In various examples, the adding operations of the block 1808 can be performed using the adders 1360L and 1360R (FIGS. 13, 14) or the adders 1560L and 1560R (FIGS. 15, 16). In some examples, the left and right output binaural signals are directed to audio-rendering devices, such as the above-mentioned XR / AR / VR headsets or earbuds of the device 1700 (FIG. 17), where left and right output binaural signals are rendered to produce the corresponding audible sounds.
[0093] According to an example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-18, provided is an audio system for generating binaural audio, the audio system comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the audio system at least to: transform a set of first audio signals corresponding to an audio scene into a set of first audio objects; decorrelate the set of first audio objects to generate a plurality of first decorrelated audio objects; apply a first head related transform filter (HRTF) to each of the first decorrelated audio objects to generate a respective left audio component and a respective right audio component; add the respective left audio components to generate a left output binaural signal; and add the respective right audio components to generate a right output binaural signal.
[0094] According to another example embodiment disclosed above, e.g., in the summary section and / or in reference to any one or any combination of some or all of FIGS. 1-18, provided is a method of generating binaural audio, the method comprising: transforming (e.g., 1802, FIG 18) a set of first audio signals corresponding to an audio scene (e.g., 1302, 1402, 1502, 1602; FIGS. 13-16) into a set of first audio objects (e.g., 1322, 1522, 1523; FIGS. 13-16); decorrelating (e.g., 1804, FIG. 18) the set of first audio objects to generate a plurality of first decorrelated audio objects (e.g., 912, FIGS. 13-16); applying (e.g., 1806, FIG. 18) a first head related transform filter (HRTF) to each of the first decorrelated audio objects to generate a respective left audio component (e.g., 1352L, 1452L, 1542LR, 1642L; FIGS. 13-16) and a respective right audio component (e.g., 1352R, 1452R, 1542R, 1642R; FIGS. 13-16); adding (e.g., 1808) the respective left audio components to generate aleft output binaural signal (e.g., 1362L, 1462L, 1562L, 1662L; FIGS. 13-16); and adding (e.g., 1808) the respective right audio components to generate a right output binaural signal (e.g., 1362R, 1462R, 1562R, 1662R; FIGS. 13-16).
[0095] In some embodiments of the above method, the set of first audio signals includes a plurality of audio channels (1302i-1302n, FIG. 13) corresponding to the audio scene.
[0096] In some embodiments of any of the above methods, the set of first audio signals includes a left input binaural signal and a right input binaural signal (1402L and 1402R, FIG. 14) corresponding to the audio scene.
[0097] In some embodiments of any of the above methods, a number of audio objects in the set of first audio objects is greater than a number of signals in the set of first audio signals (e.g., N>n, FIG. 13).
[0098] In some embodiments of any of the above methods, the decorrelating (e.g. 1804, FIG. 18) comprises: applying an individual one of the first audio objects (e.g., 702, FIG. 7; 1322i, FIGS. 13-14; 1522i, 1523i, FIGS. 15-16) to an all-pass filter (e.g., 710, FIG. 7) to generate a respective first filtered signal (e.g., 712, FIG. 7); applying the respective first filtered signal to a plurality of comb filters (e.g., 720i, FIG. 7) to generate a plurality of respective second filtered signals (e.g., yc[n], Eq. (10); yci, FIG. 7); and summing the plurality of respective second filtered signals to generate a respective summed signal (e.g., 732, FIG. 7; 912, FIG. 9).
[0099] In some embodiments of any of the above methods, the all-pass filter comprises a Schroeder all-pass section.
[0100] In some embodiments of any of the above methods, different ones of the comb filters (e.g., 720i, FIG. 7) have different respective delay parameters selected based on different respective prime numbers (e.g., Table 1).
[0101] In some embodiments of any of the above methods, each of the different respective prime numbers is greater than 100 (or 200, or 300, or 500) (e.g., Table 1).
[0102] In some embodiments of any of the above methods, the respective summed signal represents a respective one of the first decorrelated audio objects.
[0103] The method of claim 5, further comprising generating a respective one of the first decorrelated audio objects using a weighted sum of the respective summed signal and the individual one of the first audio objects (e.g., 912, FIG. 9).
[0104] In some embodiments of any of the above methods, the method further comprises performing headtracking via rotation of the first HRTF (e.g., 501, 502; FIG. 5B).
[0105] In some embodiments of any of the above methods, the applying comprises: splitting (e.g., 1110, FIG. 11) an individual one (e.g., 1102, FIG. 11; 912;, FIGS. 13-16) of the first decorrelated audio objects into a respective low-frequency audio object (e.g., 1112, FIG. 11) and a respective high-frequency audio object (e.g., 1114, FIG. 11); applying interaural-time-difference (ITD) processing (e.g., 1120, FIG. 11) and interaural-level-difference (ILD) processing (e.g., 1130i , FIG. 11) to the respective low-frequency audio object; and applying ILD processing (e.g., 11302, FIG. 11) but no ITD processing to the respective high-frequency audio object.
[0106] In some embodiments of any of the above methods, the set of first audio signals includes a left input binaural signal (e.g., 1402L, FIG. 14) and a right input binaural signal (e.g., 1402R, FIG. 14) corresponding to the audio scene; and wherein the first HRTF is normalized (e.g., 1420, 1450; FIG. 14) with respect to an HRTF corresponding to a reference head-rotation angle.
[0107] In some embodiments of any of the above methods, the method further comprises: with a low-pass filter (e.g., 1510, FIG. 15), splitting a set of input audio signals representing the audio scene into the set of first audio signals (e.g., 1512i, FIG. 15) and a set of second audio signals (e.g., 1513i, FIG. 15), wherein the first audio signals have frequencies (LF, FIG. 15) that are lower than a cutoff frequency of the low-pass filter; and wherein the second audio signals have frequencies (HF, FIG. 15) that are higher than the cutoff frequency.
[0108] In some embodiments of any of the above methods, the method further comprises: transforming the set of second audio signals (e.g., 1513;, FIG. 15) into a set of second audio objects (e.g., 1523i, FIG. 15); decorrelating (e.g., 900N-H, FIG. 15) the set of second audio objects to generate a plurality of second decorrelated audio objects; and applying a second HRTF (e.g., 1550N+i, FIG. 15) to each of the second decorrelated audio objects to generate a respective left audio component (e.g., 1552Li, FIG. 15) and a respective right audio component (e.g., 1552Ri, FIG. 15).
[0109] In some embodiments of any of the above methods, the method further comprises performing headtracking via rotation of the first HRTF and the second HRTF (e.g., 501, 502; FIG. 5B).
[0110] In some embodiments of any of the above methods, said applying the first HRTF comprises applying interaural-time-difference (ITD) processing (e.g., 1120L, 1120R; FIG. 15) and interaural-level-difference (ILD) processing (e.g., 1130IL, 1130IR, FIG. 15) to an individual one of the first decorrelated audio objects; and wherein said applying the second HRTF comprises applying ILD processing (e.g., 11302L, 11302R, FIG. 15) but no ITD processing to an individual one of the second decorrelated audio objects.
[0111] In some embodiments of any of the above methods, the set of input audio signals includes a left input binaural signal (e.g., 1602L, FIG. 16) and a right input binaural signal (e.g., 1602R, FIG. 16) corresponding to the audio scene; wherein the first HRTF is normalized (e.g., 1640i, FIG. 16) with respect to an HRTF corresponding to a reference head-rotation angle; and wherein the second HRTF is normalized (e.g., 1650i, FIG. 16) with respect to the HRTF corresponding to the reference head-rotation angle.
[0112] A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising any one of the above methods.
[0113] With regard to the processes, systems, methods, heuristics, etc. described herein, it should be understood that, although the steps of such processes, etc. have been described as occurring according to a certain ordered sequence, such processes could be practiced with the described steps performed in an order other than the order described herein. It further should be understood that certain steps could be performed simultaneously, that other steps could be added, or that certain steps described herein could be omitted. In other words, the descriptions of processes herein are provided for the purpose of illustrating certain embodiments and should in no way be construed so as to limit the claims.
[0114] Accordingly, it is to be understood that the above description is intended to be illustrative and not restrictive. Many embodiments and applications other than the examples provided would be apparent upon reading the above description. The scope should be determined, not with reference tothe above description, but should instead be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled. It is anticipated and intended that future developments will occur in the technologies discussed herein, and that the disclosed systems and methods will be incorporated into such future embodiments. In sum, it should be understood that the application is capable of modification and variation.
[0115] All terms used in the claims are intended to be given their broadest reasonable constructions and their ordinary meanings as understood by those knowledgeable in the technologies described herein unless an explicit indication to the contrary is made herein. In particular, use of the singular articles such as “a,” “the,” “said,” etc. should be read to recite one or more of the indicated elements unless a claim recites an explicit limitation to the contrary.
[0116] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments incorporate more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in fewer than all features of a single disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as a separately claimed subject matter.
[0117] While this disclosure includes references to illustrative embodiments, this specification is not intended to be construed in a limiting sense. Various modifications of the described embodiments, as well as other embodiments within the scope of the disclosure, which are apparent to persons skilled in the art to which the disclosure pertains are deemed to lie within the principle and scope of the disclosure, e.g., as expressed in the following claims.
[0118] Some embodiments may be implemented as circuit-based processes, including possible implementation on a single integrated circuit.
[0119] Some embodiments can be embodied in the form of methods and apparatuses for practicing those methods. Some embodiments can also be embodied in the form of program coderecorded in tangible media, such as magnetic recording media, optical recording media, solid state memory, floppy diskettes, CD-ROMs, hard drives, or any other non-transitory machine-readable storage medium, wherein, when the program code is loaded into and executed by a machine, such as a computer, the machine becomes an apparatus for practicing the patented invention(s). Some embodiments can also be embodied in the form of program code, for example, stored in a non- transitory machine-readable storage medium including being loaded into and / or executed by a machine, wherein, when the program code is loaded into and executed by a machine, such as a computer or a processor, the machine becomes an apparatus for practicing the patented invention(s). When implemented on a general-purpose processor, the program code segments combine with the processor to provide a unique device that operates analogously to specific logic circuits.
[0120] Unless explicitly stated otherwise, each numerical value and range should be interpreted as being approximate as if the word “about” or “approximately” preceded the value or range.
[0121] The use of figure numbers and / or figure reference labels in the claims is intended to identify one or more possible embodiments of the claimed subject matter in order to facilitate the interpretation of the claims. Such use is not to be construed as necessarily limiting the scope of those claims to the embodiments shown in the corresponding figures.
[0122] Although the elements in the following method claims, if any, are recited in a particular sequence with corresponding labeling, unless the claim recitations otherwise imply a particular sequence for implementing some or all of those elements, those elements are not necessarily intended to be limited to being implemented in that particular sequence.
[0123] Reference herein to “one embodiment” or “an embodiment” means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the disclosure. The appearances of the phrase “in one embodiment” in various places in the specification are not necessarily all referring to the same embodiment, nor are separate or alternative embodiments necessarily mutually exclusive of other embodiments. The same applies to the term “implementation.”
[0124] Unless otherwise specified herein, the use of the ordinal adjectives “first,” “second,” “third,” etc., to refer to an object of a plurality of like objects merely indicates that different instances of such like objects are being referred to, and is not intended to imply that the like objectsso referred-to have to be in a corresponding order or sequence, either temporally, spatially, in ranking, or in any other manner.
[0125] Unless otherwise specified herein, in addition to its plain meaning, the conjunction “if’ may also or alternatively be construed to mean “when” or “upon” or “in response to determining” or “in response to detecting,” which construal may depend on the corresponding specific context. For example, the phrase “if it is determined” or “if [a stated condition] is detected” may be construed to mean “upon determining” or “in response to determining” or “upon detecting [the stated condition or event]” or “in response to detecting [the stated condition or event].”
[0126] Also, for purposes of this description, the terms “couple,” “coupling,” “coupled,” “connect,” “connecting,” or “connected” refer to any manner known in the art or later developed in which energy is allowed to be transferred between two or more elements, and the interposition of one or more additional elements is contemplated, although not required. Conversely, the terms “directly coupled,” “directly connected,” etc., imply the absence of such additional elements.
[0127] As used herein in reference to an element and a standard, the term compatible means that the element communicates with other elements in a manner wholly or partially specified by the standard and would be recognized by other elements as sufficiently capable of communicating with the other elements in the manner specified by the standard. The compatible element does not need to operate internally in a manner specified by the standard.
[0128] The functions of the various elements shown in the figures, including any functional blocks labeled as “processors” and / or “controllers,” may be provided through the use of dedicated hardware as well as hardware capable of executing software in association with appropriate software. When provided by a processor, the functions may be provided by a single dedicated processor, by a single shared processor, or by a plurality of individual processors, some of which may be shared. Moreover, explicit use of the term “processor” or “controller” should not be construed to refer exclusively to hardware capable of executing software, and may implicitly include, without limitation, digital signal processor (DSP) hardware, network processor, application specific integrated circuit (ASIC), field programmable gate array (FPGA), read only memory (ROM) for storing software, random access memory (RAM), and nonvolatile storage. Other hardware, conventional and / or custom, may also be included. Similarly, any switches shown in thefigures are conceptual only. Their function may be carried out through the operation of program logic, through dedicated logic, through the interaction of program control and dedicated logic, or even manually, the particular technique being selectable by the implementer as more specifically understood from the context.
[0129] As used in this application, the terms “circuit,” “circuitry” may refer to one or more or all of the following: (a) hardware-only circuit implementations (such as implementations in only analog and / or digital circuitry); (b) combinations of hardware circuits and software, such as (as applicable): (i) a combination of analog and / or digital hardware circuit(s) with software / firmware and (ii) any portions of hardware processor(s) with software (including digital signal processor(s)), software, and memory(ies) that work together to cause an apparatus, such as a mobile phone or server, to perform various functions); and (c) hardware circuit(s) and or processor(s), such as a microprocessor(s) or a portion of a microprocessor(s), that requires software (e.g., firmware) for operation, but the software may not be present when it is not needed for operation.” This definition of circuitry applies to all uses of this term in this application, including in any claims. As a further example, as used in this application, the term circuitry also covers an implementation of merely a hardware circuit or processor (or multiple processors) or portion of a hardware circuit or processor and its (or their) accompanying software and / or firmware. The term circuitry also covers, for example and if applicable to the particular claim element, a baseband integrated circuit or processor integrated circuit for a mobile device or a similar integrated circuit in server, a cellular network device, or other computing or network device.
[0130] It should be appreciated by those of ordinary skill in the art that any block diagrams herein represent conceptual views of illustrative circuitry embodying the principles of the disclosure. Similarly, it will be appreciated that any flow charts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and so executed by a computer or processor, whether or not such computer or processor is explicitly shown.
[0131] “BRIEF SUMMARY OF SOME SPECIFIC EMBODIMENTS” in this specification is intended to introduce some example embodiments, with additional embodiments being described in “DETAILED DESCRIPTION” and / or in reference to one or more drawings. “BRIEF SUMMARYOF SOME SPECIFIC EMBODIMENTS” is not intended to identify essential elements or features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter.
Claims
CLAIMSWhat is claimed is:
1. A method of generating binaural audio, the method comprising: transforming a set of first audio signals corresponding to an audio scene into a set of first audio objects; decorrelating the set of first audio objects to generate a plurality of first decorrelated audio objects; applying a first head related transform filter (HRTF) to each of the first decorrelated audio objects to generate a respective left audio component and a respective right audio component; adding the respective left audio components to generate a left output binaural signal; and adding the respective right audio components to generate a right output binaural signal.
2. The method of claim 1, wherein the set of first audio signals includes a plurality of audio channels corresponding to the audio scene.
3. The method of claim 1, wherein the set of first audio signals includes a left input binaural signal and a right input binaural signal corresponding to the audio scene.
4. The method of any one of claims 1-3, wherein a number of audio objects in the set of first audio objects is greater than a number of signals in the set of first audio signals.
5. The method of any one of claims 1-4, wherein the decorrelating comprises: applying an individual one of the first audio objects to an all-pass filter to generate a respective first filtered signal; applying the respective first filtered signal to a plurality of comb filters to generate a plurality of respective second filtered signals; and summing the plurality of respective second filtered signals to generate a respective summed signal.
6. The method of claim 5, wherein the all-pass filter comprises a Schroeder all-pass section.
7. The method of claim 5, wherein different ones of the comb filters have different respective delay parameters selected based on different respective prime numbers.
8. The method of claim 7, wherein each of the different respective prime numbers is greater than 100 (or 200, or 300, or 500).
9. The method of claim 5, wherein the respective summed signal represents a respective one of the first decorrelated audio objects.
10. The method of claim 5, further comprising generating a respective one of the first decorrelated audio objects using a weighted sum of the respective summed signal and the individual one of the first audio objects.
11. The method of any one of claims 1-10, further comprising performing headtracking via rotation of the first HRTF.
12. The method of any one of claims 1-11, wherein the applying comprises: splitting an individual one of the first decorrelated audio objects into a respective low- frequency audio object and a respective high-frequency audio object; applying interaural-time-difference (ITD) processing and interaural-level-difference (ILD) processing to the respective low-frequency audio object; and applying ILD processing but no ITD processing to the respective high-frequency audio object.
13. The method of claim 12, wherein the set of first audio signals includes a left input binaural signal and a right input binaural signal corresponding to the audio scene; and wherein the first HRTF is normalized with respect to an HRTF corresponding to a reference head-rotation angle.
14. The method of any one of claims 1-13, further comprising:with a low-pass filter, splitting a set of input audio signals representing the audio scene into the set of first audio signals and a set of second audio signals, wherein the first audio signals have frequencies that are lower than a cutoff frequency of the low-pass filter; and wherein the second audio signals have frequencies that are higher than the cutoff frequency.
15. The method of claim 14, further comprising: transforming the set of second audio signals into a set of second audio objects; decorrelating the set of second audio objects to generate a plurality of second decorrelated audio objects; and applying a second HRTF to each of the second decorrelated audio objects to generate a respective left audio component and a respective right audio component.
16. The method of claim 15, further comprising performing headtracking via rotation of the first HRTF and the second HRTF.
17. The method of claim 15, wherein said applying the first HRTF comprises applying interaural-time-difference (ITD) processing and interaural-level-difference (ILD) processing to an individual one of the first decorrelated audio objects; and wherein said applying the second HRTF comprises applying ILD processing but no ITD processing to an individual one of the second decorrelated audio objects.
18. The method of claim 17, wherein the set of input audio signals includes a left input binaural signal and a right input binaural signal corresponding to the audio scene; wherein the first HRTF is normalized with respect to an HRTF corresponding to a reference head-rotation angle; and wherein the second HRTF is normalized with respect to the HRTF corresponding to the reference head-rotation angle.
19. A non-transitory computer-readable medium storing instructions that, when executed by an electronic processor, cause the electronic processor to perform operations comprising the method of any one of claims 1-18.
20. A system for generating binaural audio, the system comprising: at least one processor; and at least one memory including program code, wherein the at least one memory and the program code are configured to, with the at least one processor, cause the system at least to perform operations comprising the method of any one of claims 1-18.
Citation Information
Patent Citations
Spatial audio encoding and reproduction of diffuse sound
US20120082319A1
Apparatus and method for processing stereo signals for reproduction in cars to achieve individual three-dimensional sound by frontal loudspeakers
US20180014138A1
Headtracking adjusted binaural audio
WO2023059838A1